> For the complete documentation index, see [llms.txt](https://metalbear.com/mirrord/docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://metalbear.com/mirrord/docs/sharing-the-cluster/db-branching/s3.md).

# S3

Spin up an isolated clone of your remote S3 bucket with mirrord

This page covers branching for S3 object storage buckets. For the general concepts, the full list of config fields, and how a session behaves, see the [DB Branching overview](/mirrord/docs/sharing-the-cluster/db-branching.md).

Unlike the engines mirrord runs as a database server, an S3 branch is **not a pod in your cluster**: the operator has the storage provider create a new bucket in your own cloud account, clones the source bucket's settings and (optionally) its objects into it, and points your application's bucket environment variable at the clone. Your code keeps reading the same variable and never learns it is writing to a throwaway bucket.

{% hint style="info" %}
S3 branching requires operator `3.208.0`, mirrord CLI `3.252.0`, and operator Helm chart `3.208.0` with the `operator.s3Branching` value set to `true`.

If the operator doesn't support S3 branching, the session fails immediately rather than timing out: an older operator reports `mirrord operator <version> does not support feature S3 branching`, and one where the Helm value is off reports `feature S3 branching is not enabled on this mirrord operator`.
{% endhint %}

### Providers

The `provider` field selects which storage service hosts the branch bucket:

| `provider`        | Service                                 |
| ----------------- | --------------------------------------- |
| `"AWS"` (default) | [Amazon S3](https://aws.amazon.com/s3/) |

Everything on this page applies to every provider unless a section says otherwise; the [Permissions](#permissions) section is provider-specific, since each provider has its own way of granting the operator access.

### How It Works

1. The operator reads the name of the source bucket from the target workload, out of the environment variable named by the [`bucket` param](#naming-the-source-bucket).
2. It creates a branch bucket in the **source bucket's region**, named `mirrord-<10 random letters>-<source bucket name>`. When that would exceed the 63-character limit on bucket names, the source part is replaced by a hash of the source bucket's name.
3. It clones the source bucket's [configuration](#what-gets-cloned) and, depending on the [copy mode](#copy-modes), its objects.
4. Every environment variable your config lists for the `bucket` param is rewritten to the branch bucket's name, so your local process reads the clone. With a `value_pattern`, only the matched part of the value is swapped and the rest of the string stays intact.
5. When the branch's TTL expires or you run `mirrord db-branches destroy`, the operator empties the branch bucket and deletes it.

The clone runs with the operator's own cloud credentials and happens entirely inside the provider - objects never pass through the cluster or your machine.

### Basic Configuration

```json
{
  "feature": {
    "db_branches": [
      {
        "id": "uploads-bucket",
        "type": "s3",
        "provider": "AWS",
        "source": {
          "params": {
            "bucket": "UPLOADS_BUCKET"
          }
        },
        "copy": {
          "mode": "all",
          "objects": ["^fixtures/"]
        }
      }
    ]
  }
}
```

`provider` is optional and defaults to `"AWS"`.

#### Fields That Don't Apply

An S3 branch has no branch pod and no database server, so several fields from the [overview](/mirrord/docs/sharing-the-cluster/db-branching.md#key-fields) are not accepted or have no effect:

| Field              | Why                                                                                                                                                                                                                                           |
| ------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `version`, `image` | There is no container to run.                                                                                                                                                                                                                 |
| `location`         | The branch always lives in your cloud account, never in the cluster or on your machine.                                                                                                                                                       |
| `profile`          | [Branch config profiles](/mirrord/docs/sharing-the-cluster/db-branching.md#branch-config-profiles) configure a branch pod; there is no `s3BranchConfig` in the Helm chart. A branch that sets `profile` fails with an error naming the field. |
| `migrations`       | [Schema migrations](/mirrord/docs/sharing-the-cluster/db-branching/more/migrations.md) run against a database server.                                                                                                                         |
| `iam_auth`         | The operator uses its own cloud identity for the clone - see [Permissions](#permissions).                                                                                                                                                     |
| `name`             | A bucket is not a server hosting several databases.                                                                                                                                                                                           |

The usual `id`, `ttl_secs`/`ttl_mins` and `creation_timeout_secs` all work as they do for any other branch. Since cloning a large bucket takes time, raise `creation_timeout_secs` (default 60) when the session fails with a timeout while the branch is still being prepared.

### Naming the Source Bucket

The single connection param an S3 branch takes is `bucket` - the environment variable on the target that holds the source bucket's name. The field is called `source`, but `connection` is accepted as an alias, so configs that mirror the other engines work too:

```json
{
  "source": {
    "params": {
      "bucket": "UPLOADS_BUCKET"
    }
  }
}
```

The `type` field selects how the variable is resolved on the target, exactly as for the other engines, and is auto-detected when omitted:

```json
{
  "source": {
    "type": "env_from",
    "params": {
      "bucket": "UPLOADS_BUCKET"
    }
  }
}
```

The value is as flexible as any other engine's param - see [Connection Modes](/mirrord/docs/sharing-the-cluster/db-branching/more/connection.md) for the full list of sources. For example, a Kubernetes Secret:

```json
{
  "source": {
    "params": {
      "bucket": {
        "secret": "app-config",
        "key": "uploads-bucket",
        "env_var_name": "UPLOADS_BUCKET"
      }
    }
  }
}
```

a literal value:

```json
{
  "source": {
    "params": {
      "bucket": { "env_var_name": "UPLOADS_BUCKET", "value": "client-uploads" }
    }
  }
}
```

or a regex picking the bucket name out of a larger variable, so only that part of the value is swapped:

```json
{
  "source": {
    "params": {
      "bucket": {
        "env_var_name": "S3_URI",
        "value_pattern": "^s3://([^/]+)/"
      }
    }
  }
}
```

An array of sources works as well: the **first** entry locates the source bucket, and **all** entries are rewritten to the branch bucket during the session.

{% hint style="warning" %}
A `secret`, `configmap`, `gcp_secret_manager`, or `aws_secrets_manager` source must set `env_var_name`. There is no branch pod whose environment could carry the value implicitly, so mirrord needs to know which variable to fill with the branch bucket's name. A config that leaves it out is rejected before the session starts.

`bucket` is also the **only** param an S3 branch accepts - `host`, `port`, `user`, `password`, `database`, or any other extra param is a config error.
{% endhint %}

### Copy Modes

| Mode                | What gets cloned                                    | Best for                                                                 |
| ------------------- | --------------------------------------------------- | ------------------------------------------------------------------------ |
| `"empty"` (default) | The bucket and its settings, with no objects        | Apps that write their own fixtures, or source buckets too large to clone |
| `"all"`             | The settings plus the objects selected by `objects` | Reproducing a production-like bucket for debugging                       |

With `"mode": "all"`, the optional `objects` field is a list of regular expressions matched against object keys. An object is copied when it matches **any** of them; omitting the field (or passing an empty list) copies every object.

```json
{
  "copy": {
    "mode": "all",
    "objects": ["^fixtures/", "\\.json$"]
  }
}
```

Patterns are not anchored automatically, so `^` and `$` are yours to add. An invalid expression fails the branch with an error naming the pattern.

{% hint style="warning" %}
The copy runs inside the provider, so a wide `"all"` costs provider-side copy time (and request charges) rather than local bandwidth - but a bucket with many objects still takes a while to clone, and the session waits for it. Narrow the copy with `objects` whenever you can.
{% endhint %}

### What Gets Cloned

Bucket settings are cloned **best-effort**: a setting the operator cannot read (a missing permission) or that the branch bucket rejects is skipped rather than failing the branch. The branch reports what it did in its `Ready` condition - how many objects were copied, how many were skipped, and which settings were left behind - so a user whose objects or settings are missing can tell why:

```bash
kubectl get branchdatabase <branch-name> -o jsonpath='{.status.conditions[?(@.type=="Ready")].message}'
```

| Cloned                                                                      | Left out                                                             |
| --------------------------------------------------------------------------- | -------------------------------------------------------------------- |
| Object ownership controls                                                   | Event notifications and EventBridge                                  |
| Public access block                                                         | Replication                                                          |
| Bucket tags                                                                 | Lifecycle rules, access logging, static website hosting, Object Lock |
| Versioning (when enabled on the source)                                     | MFA delete                                                           |
| Default encryption                                                          | Per-object ACLs and storage classes                                  |
| CORS rules                                                                  |                                                                      |
| Bucket policy, with the source bucket's ARNs rewritten to the branch bucket |                                                                      |

Objects are copied at their latest version, with their metadata and tags. A suspended-versioning source leaves the branch unversioned, which behaves the same for new writes.

Because the bucket policy is carried over with the bucket name swapped, access that your source bucket grants through its resource policy keeps working on the branch. Access granted through an **identity policy** does not - see below.

### Permissions

The operator clones buckets with its own cloud credentials, not the target's. The rest of this section is specific to `"provider": "AWS"`.

#### The operator

The operator's AWS client is built from the default credentials provider chain. The easiest way to provide the credentials is IAM role assumption: assign an IAM role to the operator's service account (see [AWS's documentation](https://docs.aws.amazon.com/eks/latest/userguide/associate-service-account-role.html)), which the `sa.roleArn` setting in the [mirrord-operator Helm chart](https://github.com/metalbear-co/charts/blob/main/mirrord-operator/values.yaml) annotates for you.

Branch buckets are always named with the `mirrord-` prefix, so the permissions the operator needs on **its own** buckets can be scoped to `arn:aws:s3:::mirrord-*`, while the source buckets only need read access:

```json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "s3:ListBucket",
        "s3:GetObject",
        "s3:GetObjectTagging",
        "s3:GetBucketTagging",
        "s3:GetBucketVersioning",
        "s3:GetEncryptionConfiguration",
        "s3:GetBucketCORS",
        "s3:GetBucketPolicy",
        "s3:GetBucketOwnershipControls",
        "s3:GetBucketPublicAccessBlock"
      ],
      "Resource": [
        "arn:aws:s3:::client-uploads",
        "arn:aws:s3:::client-uploads/*"
      ]
    },
    {
      "Effect": "Allow",
      "Action": [
        "s3:CreateBucket",
        "s3:DeleteBucket",
        "s3:ListBucket",
        "s3:ListBucketVersions",
        "s3:PutObject",
        "s3:PutObjectTagging",
        "s3:DeleteObject",
        "s3:DeleteObjectVersion",
        "s3:PutBucketTagging",
        "s3:PutBucketVersioning",
        "s3:PutEncryptionConfiguration",
        "s3:PutBucketCORS",
        "s3:PutBucketPolicy",
        "s3:PutBucketOwnershipControls",
        "s3:PutBucketPublicAccessBlock"
      ],
      "Resource": [
        "arn:aws:s3:::mirrord-*",
        "arn:aws:s3:::mirrord-*/*"
      ]
    }
  ]
}
```

* The first statement covers the buckets your developers branch. Instead of listing them, you can apply the statement to all buckets in the account and limit it with conditions - for example a tag like `branchable=true` set on every bucket you want to allow.
* The second statement covers the branch buckets. `s3:ListBucket` on the source is what lets the operator locate its region and enumerate the objects to copy; `s3:ListBucketVersions` and `s3:DeleteObjectVersion` are what let it empty a versioned branch bucket before deleting it.

{% hint style="info" %}
If the source bucket uses SSE-KMS, its encryption configuration is cloned onto the branch, so the operator's role also needs `kms:Decrypt` and `kms:GenerateDataKey` on the key - and so does your application, for the branch bucket.
{% endhint %}

#### Your application

Your local application (and the target workload) reaches the branch bucket with its **own** credentials, not the operator's. If its access to the source bucket comes from an IAM identity policy rather than the bucket's policy, extend that policy to the branch buckets - again, `arn:aws:s3:::mirrord-*` and `arn:aws:s3:::mirrord-*/*` cover all of them.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://metalbear.com/mirrord/docs/sharing-the-cluster/db-branching/s3.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
