Acceldata
ADOC

Last updated: Sep 29, 2026 09:14 UTC

Same-account S3 Access Point authentication (IRSA or EKS Pod Identity)

Use this page when the EKS cluster and the Acceldata Data Plane already exist and everything is in the same AWS account: the cluster, the bucket, the Access Point, and the IAM role.

Goal: ADOC reaches S3 through an S3 Access Point. The IAM role is allowed on Access Point resources, the Access Point policy allows the role, and the bucket policy delegates or permits the Access Point request.

EKS Pod / ADOC workload
  ↓ IRSA or EKS Pod Identity
IAM Role in same AWS account
  ↓ S3 Access Point policy
S3 Access Point alias or ARN
  ↓ Bucket policy delegation / allow statement
S3 bucket and objects

Complete the shared setup once, then choose exactly one authentication option:

  • Option 1: IRSA
  • Option 2: EKS Pod Identity

Required Values

Token

Meaning

<region>

Same region as EKS, the bucket, and the Access Point

<account-id>

AWS account ID that hosts EKS, the bucket, the Access Point, and the IAM role

<cluster>

EKS cluster name

<namespace>

Live Data Plane namespace

<bucket>

S3 bucket name

<ap-name>

S3 Access Point name

<ap-alias>

Access Point alias after creation, ending with -s3alias

<ap-arn>

arn:aws:s3:<region>:<account-id>:accesspoint/<ap-name>

<prefix>

Allowed folder, for example allowed/

<iam-role>

Dedicated IAM role name for ADOC access

<oidc-host>

OIDC host without https://, from describe-cluster

<analysis-standalone-deploy>

Deployment that uses service account analysis-standalone-service

Service accounts in scope

Confirm live names with kubectl -n <namespace> get sa. Skip any service account that does not exist. Skip any deployment that is already failing, for example CrashLoopBackOff.

  • analysis-service
  • analysis-standalone-service
  • spark-scheduler
  • torch-monitors
  • analysis-sql-service
  • spark-history-server
  • acceldata-dataplane-spark

Shared setup

1. Create the S3 bucket

  • In S3, create bucket <bucket> in <region>.
  • Keep Block Public Access enabled.
  • Upload sample files under <prefix>, for example allowed/sample.csv.

2. Create the Access Point

  • Open bucket <bucket> and create Access Point <ap-name>.
  • Choose the same bucket <bucket>.
  • Set network to Internet or the VPC used by the cluster.
  • Keep Block Public Access enabled.
  • Copy the Access Point ARN and alias after creation.

Warning

Do not put a prefix condition on s3:ListBucket. ADOC Test Connection uses HeadBucket and does not send a prefix, so a prefix-restricted list policy can return 403 even when object reads work.

3. Create the IAM role with Access Point-only permissions

Create IAM role <iam-role>. The trust policy depends on the authentication option you choose later, but the S3 permissions are the same for both options.

Attach this inline policy. Use Access Point ARNs only. Do not add arn:aws:s3:::<bucket> to the role if you want true Access Point-only IAM.

Json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "HeadAndListAccessPoint",
      "Effect": "Allow",
      "Action": ["s3:ListBucket", "s3:GetBucketLocation"],
      "Resource": "arn:aws:s3:<region>:<account-id>:accesspoint/<ap-name>"
    },
    {
      "Sid": "ReadViaAccessPoint",
      "Effect": "Allow",
      "Action": ["s3:GetObject", "s3:GetObjectVersion"],
      "Resource": "arn:aws:s3:<region>:<account-id>:accesspoint/<ap-name>/object/<prefix>*"
    }
  ]
}

Note

This inline policy is Access Point-only. Keeping arn:aws:s3:::<bucket> out of the role means the bucket policy in step 5 is mandatory (see step 5). If you later add arn:aws:s3:::<bucket> here, step 5's bucket policy becomes optional for same-account access.

4. Add the Access Point policy

On Access Point <ap-name>, set a policy that allows only <iam-role>:

Json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "AllowListBucketViaAp",
      "Effect": "Allow",
      "Principal": { "AWS": "arn:aws:iam::<account-id>:role/<iam-role>" },
      "Action": "s3:ListBucket",
      "Resource": "arn:aws:s3:<region>:<account-id>:accesspoint/<ap-name>"
    },
    {
      "Sid": "AllowGetObjectViaAp",
      "Effect": "Allow",
      "Principal": { "AWS": "arn:aws:iam::<account-id>:role/<iam-role>" },
      "Action": ["s3:GetObject", "s3:GetObjectVersion"],
      "Resource": "arn:aws:s3:<region>:<account-id>:accesspoint/<ap-name>/object/<prefix>*"
    }
  ]
}

5. Add the bucket policy

Why this is required

In step 3 the role's inline policy names only the Access Point ARN — it does not name arn:aws:s3:::<bucket>. When a request comes through the Access Point, S3 evaluates three layers and all three must allow it:

  • Caller identity policy — allowed in step 3 (on the Access Point ARN).
  • Access Point policy — allowed in step 4.
  • Bucket policy — the bucket itself must also allow the request. This is what "Access Point delegation" means. See AWS'sAccess points policies documentation.

Because the role's identity policy does not name the bucket, this third layer cannot be satisfied by identity alone — it must be satisfied by the bucket policy. An empty bucket policy makes HeadBucket return 403 even when the Access Point policy and IAM look correct.

If you later decide to name arn:aws:s3:::<bucket> in the role's identity policy instead of staying Access Point-only, the bucket policy becomes optional for same-account access. This page keeps identity Access Point-only, so the bucket policy is required.

Which pattern to pick

If the bucket already has other statements (KMS, logging, replication), keep them and only add the statements from the chosen pattern.

Go to S3 → <bucket> → Permissions → Bucket policy. Pick one of the patterns below and preserve any unrelated existing statements.

Pattern A — Delegate to Access Points (safe default)

Mental model: the bucket steps back. It says "any request that comes through an Access Point in this account is fine; the Access Point policy and IAM decide who." Access control moves to the Access Point layer.

  • s3://<ap-alias>/... → works (step 3 IAM + step 4 Access Point policy allow it).
  • s3://<bucket>/... direct → 403. The bucket policy condition s3:DataAccessPointAccount is only set on Access Point calls; a direct bucket call does not carry that key, so the delegation does not apply, and step 3's IAM is Access Point-only, so no bucket-level allow exists either.
  • New Access Point on the same bucket later: create the Access Point, write its Access Point policy, create a new IAM role. The bucket policy stays untouched.
  • If a workload later needs direct bucket access, add arn:aws:s3:::<bucket> to that role's inline policy (Pattern A still works), or switch this bucket to Pattern C. Do not mix.
Json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DelegateToSameAccountAccessPoints",
      "Effect": "Allow",
      "Principal": "*",
      "Action": "*",
      "Resource": [
        "arn:aws:s3:::<bucket>",
        "arn:aws:s3:::<bucket>/*"
      ],
      "Condition": {
        "StringEquals": { "s3:DataAccessPointAccount": "<account-id>" }
      }
    }
  ]
}

Pattern B — Specific Access Point and role only

Mental model: the bucket names one Access Point and one role explicitly. Nothing else is trusted.

  • s3://<ap-alias>/... → works.
  • s3://<bucket>/... direct → 403 (bucket policy has no allow for direct calls; IAM in step 3 is Access Point-only).
  • New Access Point on this bucket later → denied until you add a matching statement to the bucket policy.
  • Tightest blast radius, but every future Access Point or role change forces a bucket policy edit.
Json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "AllowListViaAccessPoint",
      "Effect": "Allow",
      "Principal": { "AWS": "arn:aws:iam::<account-id>:role/<iam-role>" },
      "Action": ["s3:ListBucket", "s3:GetBucketLocation"],
      "Resource": "arn:aws:s3:::<bucket>",
      "Condition": {
        "StringEquals": { "s3:DataAccessPointArn": "arn:aws:s3:<region>:<account-id>:accesspoint/<ap-name>" }
      }
    },
    {
      "Sid": "AllowGetObjectViaAccessPoint",
      "Effect": "Allow",
      "Principal": { "AWS": "arn:aws:iam::<account-id>:role/<iam-role>" },
      "Action": ["s3:GetObject", "s3:GetObjectVersion"],
      "Resource": "arn:aws:s3:::<bucket>/*",
      "Condition": {
        "StringEquals": { "s3:DataAccessPointArn": "arn:aws:s3:<region>:<account-id>:accesspoint/<ap-name>" }
      }
    }
  ]
}

Pattern C — Hybrid bucket + Access Point

Mental model: the bucket allows the role directly. Access Point is one path; direct s3://<bucket> is another. Both work with the same role. Not Access Point-only. Use this when the same role must also access s3://<bucket> or s3a://<ap-alias>/....

  • s3://<ap-alias>/... → works.
  • s3://<bucket>/... direct → works.
  • Use only when a real workload needs direct bucket access, for example Spark event-log write, or a legacy script that cannot switch to the Access Point alias.
Json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "AllowRoleOnBucket",
      "Effect": "Allow",
      "Principal": { "AWS": "arn:aws:iam::<account-id>:role/<iam-role>" },
      "Action": [
        "s3:ListBucket",
        "s3:GetBucketLocation",
        "s3:GetObject",
        "s3:GetObjectVersion",
        "s3:PutObject",
        "s3:DeleteObject"
      ],
      "Resource": ["arn:aws:s3:::<bucket>", "arn:aws:s3:::<bucket>/*"]
    }
  ]
}

Recap

  • Pattern A is the safe default. Direct bucket calls return 403; every Access Point flow works. Cheapest to extend.
  • Pattern B is A minus future flexibility, plus an explicit role binding.
  • Pattern C is the only option when s3://<bucket> or s3a://<bucket> must keep working.

Option 1: IRSA

  • Confirm the cluster OIDC provider
Bash
aws eks describe-cluster --name <cluster> --region <region> \
  --query 'cluster.identity.oidc.issuer' --output text

If the result is empty, associate the provider:

Bash
eksctl utils associate-iam-oidc-provider --cluster <cluster> --region <region> --approve
  • Configure the trust policy

In IAM, open <iam-role> → Trust. Configure Web Identity for this cluster's OIDC provider, with audience sts.amazonaws.com. Set sub to every service account listed above.

Json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Principal": { "Federated": "arn:aws:iam::<account-id>:oidc-provider/<oidc-host>" },
      "Action": "sts:AssumeRoleWithWebIdentity",
      "Condition": {
        "StringEquals": { "<oidc-host>:aud": "sts.amazonaws.com" },
        "StringLike": {
          "<oidc-host>:sub": [
            "system:serviceaccount:<namespace>:analysis-service",
            "system:serviceaccount:<namespace>:analysis-standalone-service",
            "system:serviceaccount:<namespace>:spark-scheduler",
            "system:serviceaccount:<namespace>:torch-monitors",
            "system:serviceaccount:<namespace>:analysis-sql-service",
            "system:serviceaccount:<namespace>:spark-history-server",
            "system:serviceaccount:<namespace>:acceldata-dataplane-spark"
          ]
        }
      }
    }
  ]
}

A Pod Identity association on the same service account overrides IRSA. For IRSA-only, delete those associations first.

Bash
aws eks list-pod-identity-associations --cluster-name <cluster> --region <region> \
  --query 'associations[?namespace==`<namespace>`].[associationId,serviceAccount]' --output table
aws eks delete-pod-identity-association --cluster-name <cluster> --region <region> \
  --association-id <association-id>
  • Bind the role to service accounts
Bash
NS=<namespace>
ROLE=arn:aws:iam::<account-id>:role/<iam-role>
for sa in analysis-service analysis-standalone-service spark-scheduler torch-monitors analysis-sql-service spark-history-server acceldata-dataplane-spark
do
 kubectl -n "sa" "eks.amazonaws.com/role-arn=${ROLE}" --overwrite
done
Bash
kubectl -n <namespace> exec deploy/<analysis-standalone-deploy> -- env | \
 grep -E 'AWS_ROLE_ARN|AWS_WEB_IDENTITY_TOKEN_FILE|AWS_CONTAINER_CREDENTIALS'
  • Verify IRSA credentials in the pod

Pass:AWS_ROLE_ARN=arn:aws:iam::<account-id>:role/<iam-role> is present and AWS_CONTAINER_CREDENTIALS_FULL_URI is absent. Authentication in ADOC is AWS IAM Roles For Service Accounts.

Option 2: EKS Pod Identity

  • Check the Pod Identity agent
Bash
kubectl -n kube-system get ds eks-pod-identity-agent
  • Configure the trust policy

In IAM, open <iam-role> → Trust and add this statement. If the role already has the IRSA trust, keep the OIDC statement and add this one. Do not change the Access Point policy, bucket policy, or S3 inline policy.

Json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Principal": { "Service": "pods.eks.amazonaws.com" },
      "Action": ["sts:AssumeRole", "sts:TagSession"]
    }
  ]
}
  • Bind associations

Do not annotate the service accounts for Pod Identity.

Bash
CLUSTER=<cluster>
REGION=<region>
NS=<namespace>
ROLE=arn:aws:iam::<account-id>:role/<iam-role>
for sa in analysis-service analysis-standalone-service spark-scheduler torch-monitors analysis-sql-service spark-history-server acceldata-dataplane-spark
do
  aws eks create-pod-identity-association \
    --cluster-name "REGION" \
    --namespace "sa" \
    --role-arn "$ROLE"
done

Skip service accounts that do not exist, then restart the deployments.

Bash
kubectl -n <namespace> rollout restart deploy

Wait until the deployments are ready. Spark job pods pick this up on the next crawl.

  • Verify Pod Identity credentials in the pod
Bash
kubectl -n <namespace> exec deploy/<analysis-standalone-deploy> -- env | \
 grep -E 'AWS_ROLE_ARN|AWS_WEB_IDENTITY_TOKEN_FILE|AWS_CONTAINER_CREDENTIALS'

Pass:AWS_CONTAINER_CREDENTIALS_FULL_URI is set and AWS_ROLE_ARN is gone, because Pod Identity overrides IRSA. Authentication in ADOC is EKS Pod Identity. Use a new ADOC datasource if IRSA was already added.

Prove AWS access before configuring ADOC

Delete any leftover test pod first; old logs are not proof.

Bash
kubectl -n <namespace> delete pod ap-head --ignore-not-found
kubectl -n <namespace> run ap-head --restart=Never \
 --image=public.ecr.aws/aws-cli/aws-cli:2.17.55 \
 --overrides='{"spec":{"serviceAccountName":"analysis-standalone-service"}}' \
 --command -- aws s3api head-bucket --bucket <ap-arn> --region <region>

Wait until the pod status is Completed, then run:

Bash
kubectl -n <namespace> logs ap-head

Pass: JSON containing BucketRegion. A 403 indicates IAM, Access Point policy, or bucket policy — not ADOC. Do not exec into this pod; it has already exited. An Access Point-only role on torch-monitors may return 403 for existing SQS monitors; this is expected.

ADOC: add datasource and validate

In ADOC, add a new AWS S3 datasource. For the general steps to open the Add Data Source wizard, seeAmazon S3. Use these values for an Access Point connection.

  • Add datasource / Test Connection

In the tenant, add AWS S3 with these values:

Connection field

Value

Region

<region>

Authentication

Option 1: AWS IAM Roles For Service Accounts. Option 2: EKS Pod Identity

Bucket Name

<bucket>if using Pattern C, otherwise <ap-alias>/<ap-arn> if the alias fails. Do not use s3a:// URI.

Data Plane

This cluster's Data Plane in <namespace>

Warning

Never enter s3a://... in Bucket Name — that scheme belongs in Path Expression below. Bucket Name takes only the bare alias or ARN, since that's the string AWS's HeadBucket call uses internally. Click Test Connection; pass means connected.

  • Asset / Observability configuration

After the connection succeeds, create an asset with any name and set:

Field

Value

Path Expression

S3 Object URI e.g. s3://<bucket>/<prefix>* or

s3a://<ap-alias>/<prefix>*

File Type / delimiter

Match the files under <prefix>

Crawler schedule / notify / cadence

Optional

Then crawl and profile. GET /validate on analysis-service is a health probe; jobs are CONNECTION_VALIDATION.

Troubleshooting quick checks

Symptom

Likely Cause

Check

HeadBucket 403 during Test Connection

Bucket policy does not allow the Access Point request, or s3:prefix restricts list behavior.

Recheck Access Point policy, bucket policy Pattern A/B/C, and remove prefix conditions from s3:ListBucket.

IRSA env vars missing

Service account annotation missing, wrong trust policy, or a Pod Identity association overrides IRSA.

Run the IRSA verification command and list Pod Identity associations.

Pod Identity credential URI missing

Pod Identity Agent not installed or association missing for the service account.

Check eks-pod-identity-agent and aws eks list-pod-identity-associations.

Direct s3://<bucket> or s3a://<bucket>/... returns 403 (Pattern A or B)

IAM identity is Access Point-only, so direct bucket calls have no identity allow.

Either add arn:aws:s3:::<bucket> to the role's inline policy in step 3, or switch the bucket policy to Pattern C.