Cross-account S3 Access Point authentication (IRSA or EKS Pod Identity)
This page configures ADOC running on EKS in Account A to read data through an S3 Access Point hosted in Account B. The Data Plane already runs in Account A. The bucket and Access Point live in Account B. Pods in Account A use one dedicated IAM role, and ADOC connects to the Access Point as the datasource.
Note
Do the shared setup once, then choose exactly one authentication path: Option 1: IRSA or Option 2: EKS Pod Identity.
Required values
Token | Meaning |
<region> | Same region as EKS, the bucket, and the Access Point |
<account-a-id> | Account A that hosts EKS |
<account-b-id> | Account B that hosts the bucket and Access Point |
<cluster> | EKS cluster name in Account A |
<namespace> | Live Data Plane namespace |
<bucket> | Bucket name in Account B |
<ap-name> | Access Point name in Account B |
<ap-alias> | Access Point alias after creation, ends with -s3alias |
<ap-arn> | arn:aws:s3:<region>:<account-b-id>:accesspoint/<ap-name> |
<prefix> | Allowed folder, for example allowed/ |
<iam-role> | Dedicated IAM role in Account A. Create a new role and do not reuse a same-account Access Point role. |
<oidc-host> | OIDC host without https:// |
<analysis-standalone-deploy> | Deployment that uses service account analysis-standalone-service |
<spark-history-bucket> | Bucket in Account A used for Spark event logs |
Note
Confirm service account names with kubectl -n <namespace> get sa. There is no service account named dataplane-spark. Skip any service account that does not exist. Skip any deployment that is already failing, for example CrashLoopBackOff.
Service accounts in scope
- analysis-service
- analysis-standalone-service
- spark-scheduler
- torch-monitors
- analysis-sql-service
- spark-history-server
- acceldata-dataplane-spark
Shared setup
1. Create the bucket in Account B
- Log in to Account B.
- Create bucket <bucket> in region <region>.
- Keep Block Public Access enabled.
- Upload sample files under <prefix>, for example allowed/sample.csv.
- Leave the bucket policy empty until the IAM role in Account A exists.
2. Create the Access Point in Account B
- In S3, open bucket <bucket> and create Access Point <ap-name>.
- Choose the same bucket <bucket>.
- Set network to Internet or the VPC used by the cluster.
- Keep Block Public Access enabled.
- Copy the Access Point ARN and alias.
Warning
Do not put a prefix condition on s3:ListBucket. ADOC Test Connection uses HeadBucket and does not send a prefix.
3. Create the IAM role and inline S3 policy in Account A
Create IAM role <iam-role> in Account A. Set the trust relationship later in Option 1 or Option 2. Attach this inline S3 policy to the role:
Json
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "HeadAndList",
"Effect": "Allow",
"Action": ["s3:ListBucket", "s3:GetBucketLocation"],
"Resource": [
"arn:aws:s3:<region>:<account-b-id>:accesspoint/<ap-name>",
"arn:aws:s3:::<bucket>"
]
},
{
"Sid": "ReadObjects",
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:GetObjectVersion"],
"Resource": [
"arn:aws:s3:<region>:<account-b-id>:accesspoint/<ap-name>/object/<prefix>*",
"arn:aws:s3:::<bucket>/<prefix>*"
]
}
]
}
Cross-account access requires both the Access Point ARN and the bucket ARN on the IAM role.
4. Add the Access Point policy in Account B
On Access Point <ap-name>, allow only the Account A role:
Json
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "AllowListBucketViaAp",
"Effect": "Allow",
"Principal": { "AWS": "arn:aws:iam::<account-a-id>:role/<iam-role>" },
"Action": "s3:ListBucket",
"Resource": "arn:aws:s3:<region>:<account-b-id>:accesspoint/<ap-name>"
},
{
"Sid": "AllowGetObjectViaAp",
"Effect": "Allow",
"Principal": { "AWS": "arn:aws:iam::<account-a-id>:role/<iam-role>" },
"Action": ["s3:GetObject", "s3:GetObjectVersion"],
"Resource": "arn:aws:s3:<region>:<account-b-id>:accesspoint/<ap-name>/object/<prefix>*"
}
]
}
5. Add the bucket policy in Account B
Why this is required
In step 3 the Account A role policy names both the Access Point ARN and the bucket ARN, because this is cross-account access. When a request comes through the Access Point, S3 still evaluates three layers and all three must allow it:
- Caller identity policy — allowed in step 3 on both the Access Point ARN and the bucket ARN.
- Access Point policy — allowed in step 4 for the Account A role.
- Bucket policy — the Account B bucket itself must also allow the request. This is what "Access Point delegation" means. See AWS'sConfiguring IAM policies for using access points documentation.
For cross-account access, the bucket policy is required. An empty bucket policy returns 403 even when the Access Point policy and the Account A IAM role policy look correct.
If the bucket already has other statements, for example KMS, logging, replication, or unrelated application access, keep them and only add the statements from the chosen pattern.
Which pattern to pick
Go to S3 → <bucket> → Permissions → Bucket policy. Pick one of the patterns below and preserve any unrelated existing statements.
Pattern A — Delegate to Access Points (safe default)
Mental model: the bucket steps back. It says "any request that comes through an Access Point in Account B is fine; the Access Point policy and IAM decide who." Access control moves to the Access Point layer.
- s3://<ap-alias>/... → works when step 3 IAM and step 4 Access Point policy allow it.
- s3://<bucket>/... direct → not allowed by this delegation pattern unless another existing bucket policy or identity policy allows direct bucket access. A direct bucket call does not carry the s3:DataAccessPointAccount condition key, so this delegation does not apply.
- New Access Point on the same bucket later: create the Access Point, write its Access Point policy, and create or update the IAM role. The bucket policy stays untouched.
- If a workload later needs direct bucket access, add direct bucket permissions explicitly, or use Pattern C. Do not mix patterns unnecessarily.
Json
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "DelegateToAccountBAccessPoints",
"Effect": "Allow",
"Principal": "*",
"Action": "*",
"Resource": [
"arn:aws:s3:::<bucket>",
"arn:aws:s3:::<bucket>/*"
],
"Condition": {
"StringEquals": { "s3:DataAccessPointAccount": "<account-b-id>" }
}
}
]
}
Pattern B — Specific Access Point and Account A role only
Mental model: the bucket names one Access Point and one Account A role explicitly. Nothing else is trusted.
- s3://<ap-alias>/... → works.
- s3://<bucket>/... direct → not allowed by this pattern unless another existing bucket policy or identity policy allows direct bucket access.
- New Access Point on this bucket later → denied until you add a matching statement to the bucket policy.
- Tightest blast radius, but every future Access Point or role change forces a bucket policy edit.
Json
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "AllowListViaAccessPoint",
"Effect": "Allow",
"Principal": { "AWS": "arn:aws:iam::<account-a-id>:role/<iam-role>" },
"Action": ["s3:ListBucket", "s3:GetBucketLocation"],
"Resource": "arn:aws:s3:::<bucket>",
"Condition": {
"StringEquals": { "s3:DataAccessPointArn": "arn:aws:s3:<region>:<account-b-id>:accesspoint/<ap-name>" }
}
},
{
"Sid": "AllowGetObjectViaAccessPoint",
"Effect": "Allow",
"Principal": { "AWS": "arn:aws:iam::<account-a-id>:role/<iam-role>" },
"Action": ["s3:GetObject", "s3:GetObjectVersion"],
"Resource": "arn:aws:s3:::<bucket>/*",
"Condition": {
"StringEquals": { "s3:DataAccessPointArn": "arn:aws:s3:<region>:<account-b-id>:accesspoint/<ap-name>" }
}
}
]
}
Pattern C — Hybrid bucket + Access Point
Mental model: the bucket allows the Account A role directly. Access Point is one path; direct s3://<bucket> is another. Both work with the same role. Not Access Point-only. Use this when the same role must also access s3://<bucket> or s3a://<bucket>/....
- s3://<ap-alias>/... → works.
- s3://<bucket>/... direct → works.
- Use only when a real workload needs direct bucket access, for example Spark event-log write, or a legacy script that cannot switch to the Access Point alias.
Json
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "AllowRoleOnBucket",
"Effect": "Allow",
"Principal": { "AWS": "arn:aws:iam::<account-a-id>:role/<iam-role>" },
"Action": [
"s3:ListBucket",
"s3:GetBucketLocation",
"s3:GetObject",
"s3:GetObjectVersion"
],
"Resource": ["arn:aws:s3:::<bucket>", "arn:aws:s3:::<bucket>/*"]
}
]
}
Recap
- Pattern A is the safe default. Direct bucket calls are not granted by this pattern; Access Point flows are delegated to the Access Point policy and IAM.
- Pattern B is Pattern A with a smaller blast radius and less future flexibility because it binds to a specific Access Point and Account A role.
- Pattern C is the only option when s3://<bucket> or s3a://<bucket> must keep working.
Option 1: IRSA
6. Confirm or create the OIDC provider
Bash
aws eks describe-cluster --name <cluster> --region <region> --query 'cluster.identity.oidc.issuer' --output text
Strip https:// from the result. That value is <oidc-host>. If the result is empty, create the provider:
Bash
eksctl utils associate-iam-oidc-provider --cluster <cluster> --region <region> --approve
7. Set the IRSA trust policy
Update the trust relationship on role <iam-role>:
Json
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": { "Federated": "arn:aws:iam::<account-a-id>:oidc-provider/<oidc-host>" },
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": { "<oidc-host>:aud": "sts.amazonaws.com" },
"StringLike": {
"<oidc-host>:sub": [
"system:serviceaccount:<namespace>:analysis-service",
"system:serviceaccount:<namespace>:analysis-standalone-service",
"system:serviceaccount:<namespace>:spark-scheduler",
"system:serviceaccount:<namespace>:torch-monitors",
"system:serviceaccount:<namespace>:analysis-sql-service",
"system:serviceaccount:<namespace>:spark-history-server",
"system:serviceaccount:<namespace>:acceldata-dataplane-spark"
]
}
}
}
]
}
8. Remove Pod Identity associations for these service accounts, if any
If a service account has both Pod Identity and IRSA, Pod Identity wins. For IRSA-only, remove those associations first.
Bash
aws eks list-pod-identity-associations --cluster-name <cluster> --region <region> \
--query 'associations[?namespace==`<namespace>`].[associationId,serviceAccount,roleArn]' --output table
Delete each matching association:
Bash
aws eks delete-pod-identity-association --cluster-name <cluster> --region <region> \
--association-id <association-id>
9. Annotate the service accounts
Bash
NS=<namespace>
ROLE=arn:aws:iam::<account-a-id>:role/<iam-role>
for sa in analysis-service analysis-standalone-service spark-scheduler torch-monitors analysis-sql-service spark-history-server acceldata-dataplane-spark
do
kubectl -n "sa" eks.amazonaws.com/role-arn="$ROLE" --overwrite
done
Bash
kubectl -n <namespace> rollout restart deploy
10. Restart deployments
Wait until the deployments are ready. If needed, find the standalone deployment:
Bash
kubectl -n <namespace> get deploy | grep analysis-standalone
Bash
kubectl -n <namespace> exec deploy/<analysis-standalone-deploy> -- env | \
grep -E 'AWS_ROLE_ARN|AWS_WEB_IDENTITY_TOKEN_FILE|AWS_CONTAINER_CREDENTIALS'
11. Confirm IRSA on a pod
Expected result for IRSA:AWS_ROLE_ARN=arn:aws:iam::<account-a-id>:role/<iam-role> is present and AWS_CONTAINER_CREDENTIALS_FULL_URI is not present.
Option 2: EKS Pod Identity
6. Confirm the Pod Identity agent
Bash
kubectl -n kube-system get ds eks-pod-identity-agent
The agent must be running.
7. Set the Pod Identity trust policy
In the role, add this trust policy statement:
Json
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": { "Service": "pods.eks.amazonaws.com" },
"Action": ["sts:AssumeRole", "sts:TagSession"]
}
]
}
If you already completed IRSA setup, add this statement to the existing trust policy and keep the OIDC statement. sts:TagSession is required.
8. Recreate service account associations
List current associations:
Bash
aws eks list-pod-identity-associations --cluster-name <cluster> --region <region> \
--query 'associations[?namespace==`<namespace>`].[associationId,serviceAccount,roleArn]' --output table
If a target service account already has an association, delete it before creating the new one:
Bash
aws eks delete-pod-identity-association --cluster-name <cluster> --region <region> \
--association-id <association-id>
Create the associations:
Bash
CLUSTER=<cluster>
REGION=<region>
NS=<namespace>
ROLE=arn:aws:iam::<account-a-id>:role/<iam-role>
for sa in analysis-service analysis-standalone-service spark-scheduler torch-monitors analysis-sql-service spark-history-server acceldata-dataplane-spark
do
aws eks create-pod-identity-association --cluster-name "$CLUSTER" \
--region "NS" --service-account "$sa" \
--role-arn "$ROLE"
done
You can leave the IRSA annotation on the service accounts. When both exist, Pod Identity is used instead of IRSA.
9. Restart deployments
Bash
kubectl -n <namespace> rollout restart deploy
Wait until the deployments are ready. Spark job pods pick this up on the next crawl.
10. Confirm Pod Identity on a pod
Bash
kubectl -n <namespace> exec deploy/<analysis-standalone-deploy> -- env | \
grep -E 'AWS_ROLE_ARN|AWS_WEB_IDENTITY_TOKEN_FILE|AWS_CONTAINER_CREDENTIALS'
Expected result for Pod Identity:AWS_CONTAINER_CREDENTIALS_FULL_URI is present and AWS_ROLE_ARN is absent.
Verify access from the cluster
Run these checks after completing either Option 1 or Option 2. Delete any leftover ap-head pod first. Create a new pod each time. Read logs after the pod reaches Completed. Do not exec into it.
1. HeadBucket using the Access Point ARN
Bash
kubectl -n <namespace> delete pod ap-head --ignore-not-found
kubectl -n <namespace> run ap-head --restart=Never \
--image=public.ecr.aws/aws-cli/aws-cli:2.17.55 \
--overrides='{"spec":{"serviceAccountName":"analysis-standalone-service"}}' \
--command -- aws s3api head-bucket --bucket <ap-arn> --region <region>
kubectl -n <namespace> logs ap-head
Pass condition: the logs return JSON that includes BucketRegion.
2. HeadObject using the Access Point alias
Bash
kubectl -n <namespace> delete pod ap-head --ignore-not-found
kubectl -n <namespace> run ap-head --restart=Never \
--image=public.ecr.aws/aws-cli/aws-cli:2.17.55 \
--overrides='{"spec":{"serviceAccountName":"analysis-standalone-service"}}' \
--command -- aws s3api head-object --bucket <ap-alias> --key <prefix>sample.csv --region <region>
kubectl -n <namespace> logs ap-head
Pass condition: the logs return JSON that includes ContentLength or ETag.
Configure the ADOC datasource
In ADOC, add a new AWS S3 datasource. For the general steps to open the Add Data Source wizard, seeAmazon S3. Use these values for a cross-account Access Point connection.
Connection field | Value |
Region | <region> |
Authentication | Option 1: AWS IAM Roles For Service Accounts. Option 2: EKS Pod Identity |
Bucket Name |
|
Data Plane | This cluster's Data Plane in <namespace> |
Click Test Connection. A successful result shows the datasource connected.
Asset or Observability configuration
Field | Value |
Asset Name | Any label |
Path Expression | S3 Object URI e.g. s3a://<ap-alias>/<prefix>* |
File Type / delimiter | Match the files under <prefix> |
Crawler schedule / notify / cadence | Optional |
Run the crawl, then run the profile.
GET /validate on analysis-service is a health probe. The relevant jobs for datasource testing use type CONNECTION_VALIDATION.
Troubleshooting
- If ADOC Test Connection fails on bucket access, verify that the bucket policy in Account B allows the role from Account A in addition to the Access Point policy.
- If HeadBucket fails but HeadObject works, recheck that no prefix condition was added to s3:ListBucket.
- If IRSA is configured but pods still use container credentials, look for existing Pod Identity associations on the same service account.
- If Pod Identity is configured but credentials are missing, confirm the eks-pod-identity-agent daemonset is running and the trust policy includes sts:TagSession.
- If the alias fails in ADOC, use <ap-arn> for connection testing, then confirm alias resolution separately.

Have a suggestion?