Skip to main content

Migrating to New Object Storage

This guide explains how to move the S3-compatible storage behind IOMETE to new storage, for example new hardware, a new data center or another S3-compatible product, when the bucket names stay the same and the storage endpoint changes. It covers what to change in IOMETE, what to restart, how to cut over without losing data and how to check the result.

Bucket names must stay the same

Iceberg table metadata stores absolute paths such as s3a://lakehouse/data/.... These paths contain the bucket name but not the endpoint, which is why an endpoint change is safe. Moving data to a bucket with a different name breaks every table and is not covered by this guide. Contact IOMETE support if you need to rename a bucket.

The guide is written for IOMETE 4.0. Differences in 3.19 and earlier are called out where they apply.

Choosing an Approach​

There are two ways to point IOMETE at the new storage. Both need the same data copy, write freeze, restarts and validation.

Keep the old host name (recommended when you control DNS and certificates). Point the old storage host name at the new storage through DNS, and give the new storage a certificate that is valid for that name. IOMETE's configuration, the settings entered manually and tools outside IOMETE all stay as they are, and rolling back means pointing the name back. You skip Changing the Endpoint and the legacy job refresh. Keep a way to reach the old storage after the switch, such as its IP address or a temporary host name, for the object comparison and for rollback.

Change the endpoint. Use this when the old host name can't move to the new storage. You update the endpoint in the Helm values and in every place it was entered manually, as described in Changing the Endpoint.

How IOMETE Uses Object Storage​

Table data and Iceberg metadata files live in your buckets. The catalog, which stores a pointer to each table's current metadata file, lives in IOMETE's PostgreSQL database. Because both refer to storage by bucket name only, existing tables stay valid after the move.

PostgreSQL is not part of the migration, so the new storage must match the catalog exactly. If a table changes after the final copy, the catalog points to a metadata file that is missing from the new storage, and that table fails to open. The cutover procedure prevents this.

Besides table data, IOMETE keeps the following in the lakehouse bucket. Copy the whole bucket so that nothing is left behind.

Location in the bucketContentRecommendation
data/Table data and Iceberg metadata filesRequired.
iomete-assets/sql-editor/worksheets/Saved SQL Editor worksheets. The database stores only their location.Required.
iomete-assets/spark-history/Spark event logs shown in the Spark History ServerRecommended. Without it, the history of past Spark applications is lost.
ranger/auditData access audit logsRecommended, especially for audit or compliance requirements.
Job run log location, if job run logging to S3 is enabledLogs of orchestrated job runsRecommended.
iomete-assets/sql-results/, iomete-assets/arrow-fetchTemporary query resultsOptional. They are recreated as needed.

Also copy any other bucket on the same storage that IOMETE uses, such as buckets of additional Spark catalogs or of Storage Configs. Notebooks in Jupyter containers are stored on Kubernetes volumes, not in S3, and are not affected.

Changing the Endpoint​

Prepare the changes in this section in advance, but apply them only in the Switch step of the cutover, after the final copy. Applied earlier, they let workloads write to the new storage before it holds a complete copy.

Helm Values​

The storage endpoint is set once, in the Helm values of the release that owns the storage (see Installing IOMETE). During helm upgrade, IOMETE writes it into the configuration of its services and of every Spark workload, including the Hadoop configuration Secret it creates in each connected namespace. Change the endpoint there and nowhere else. Secrets and ConfigMaps generated by IOMETE are overwritten on the next upgrade, so editing them manually does not work.

values.yaml
storage:
bucketName: "lakehouse" # unchanged
type: "s3_compatible"
s3CompatibleSettings:
endpoint: "https://<new-endpoint-host>:<port>"
accessKey: "<unchanged>"
secretKey: "<unchanged>"

In the switch step, apply the change with the chart and chart version you already run, so the endpoint change does not also become a version upgrade. helm list -n <namespace> shows both in its CHART column.

helm upgrade <release-name> <chart> \
--namespace <namespace> \
--version <current-chart-version> \
-f values.yaml
IOMETE 4.0 with several data planes

Each data plane has its own storage. The iomete-data-plane-enterprise release holds the storage of the control plane and its default data plane. An additional data plane holds its own storage in its own release of the iomete-lakehouse-runtime chart. Update the release whose storage you are moving, with the chart it was installed from.

IOMETE 3.19.0 and earlier

The storage settings sit under minioSettings (type minio) or dellEcsSettings (type dell_ecs) instead of s3CompatibleSettings. Keep the block you use today and change only endpoint. Newer versions still accept these keys.

Settings Entered Manually​

The Helm change covers everything IOMETE configures by default. The endpoint may also have been entered manually in the following places, which IOMETE does not update. During preparation, open each one and note every value that contains the old host name. Replace them in the switch step.

WhereWhat to check
Spark CatalogsEach catalog's custom S3 credentials and additional properties.
Global Spark SettingsEvery Spark property.
Compute clustersEach cluster's Spark configuration.
Spark jobs and streaming jobsEach job's Spark configuration.
Storage ConfigsThe endpoint URL of each storage config on the same storage.
Job run log settingsservices.jobOrchestrator.s3Logging.endpoint in the Helm values, if set. When empty, it follows the main storage setting.
Your own secretsIOMETE secrets, Kubernetes Secrets or HashiCorp Vault entries that contain the endpoint.
Tools outside IOMETEdbt, Airflow, BI tools, external Spark or Trino clusters, ingestion pipelines, and log systems such as Loki that use the same storage.

In Spark settings, the endpoint usually appears in these keys, but check all values for the old host name:

  • spark.hadoop.fs.s3a.endpoint
  • spark.sql.catalog.<catalog-name>.s3.endpoint
  • spark.hadoop.fs.s3a.bucket.<bucket-name>.endpoint (see Accessing Specific Buckets)

Finding Remaining References​

These commands search the Kubernetes cluster for the old host name and print only where it appears, not the values. Use the host name alone (for example old-s3.example.com), without https:// or a port. They need jq 1.6 or newer and permission to read Secrets in all namespaces.

Secrets and ConfigMaps
OLD_HOST="<old-endpoint-host>"

kubectl get secrets -A -o json | jq -r --arg h "$OLD_HOST" '.items[]
| "\(.metadata.namespace)/\(.metadata.name)" as $n
| (.data // {}) | to_entries[]
| select(.value | @base64d | contains($h))
| "secret \($n) key: \(.key)"'

kubectl get configmaps -A -o json | jq -r --arg h "$OLD_HOST" '.items[]
| "\(.metadata.namespace)/\(.metadata.name)" as $n
| (.data // {}) | to_entries[]
| select(.value | contains($h))
| "configmap \($n) key: \(.key)"'
Legacy scheduled Spark jobs
kubectl get scheduledsparkapplications -A -o json | jq -r --arg h "$OLD_HOST" '.items[]
| select(.spec.template.sparkConf // {} | tostring | contains($h))
| "\(.metadata.namespace)/\(.metadata.name)"'

Before the switch, entries generated by IOMETE (such as hadoop-config) and every legacy scheduled job appear in the results. After the switch and the restarts, all three commands should return nothing. The name printed for a scheduled job is its IOMETE job ID.

Certificates and Network Access​

  • TLS. If the new endpoint uses HTTPS, its certificate must include the new host name. If a different certificate authority issues it, add that authority to the Java truststore and the CA certificate IOMETE uses (javaTrustStore in the Helm values) before the switch.
  • Network. The new endpoint must be reachable from every namespace where IOMETE services and Spark workloads run, and in IOMETE 4.0 also from the control plane, which reads SQL Editor results directly from each data plane's storage. Arrow Flight clients and external engines that use the IOMETE Iceberg REST catalog also read data straight from storage, so the endpoint must be reachable from their networks too. They receive the new endpoint automatically.

Restarting Services and Workloads​

IOMETE services read the storage settings when they start, so restart them after the switch. When you keep the old host name, the restart also clears cached DNS addresses and open connections to the old storage. In IOMETE 4.0, upgrading the iomete-data-plane-enterprise release restarts iom-core, iom-sql, iom-cluster and iom-rest-catalog when the storage values change, but not the other services, so restart all of them. In 3.19.0 and earlier, no service restarts when only the endpoint changes.

  1. Restart all IOMETE services in the namespace of the release you upgraded:

    kubectl rollout restart deployment,statefulset -n <namespace> -l app.kubernetes.io/instance=<release-name>
    kubectl get pods -n <namespace> -w

    If you moved the storage of an additional data plane, the control plane picks up the new endpoint within two minutes. Restart iom-cluster on the control plane if you don't want to wait.

  2. Restart running workloads from the console: compute clusters, Spark Connect clusters, Jupyter containers, and streaming or other long-running Spark jobs. They keep the endpoint they started with until restarted.

  3. Refresh legacy scheduled Spark jobs (endpoint change only). Scheduled jobs on the LEGACY deployment flow run from a copy of the Spark configuration stored in Kubernetes when the job was last saved, so they keep the old endpoint. Jobs on the Priority-Based flow and manual runs are not affected. To refresh a legacy job, suspend and resume it, or open and save it without changes, through the console or API. Suspending and resuming jobs through IOMETE during the cutover does this automatically.

Copying the Data​

S3-Level Copy​

Copy through the S3 API with any tool that can sync two S3 endpoints. This works between any S3-compatible products and does not depend on how either side stores data on disk.

  • Lakehouse Backup Job: runs on IOMETE, suited to the initial bulk copy while the platform is running.
  • rclone sync old:<bucket> new:<bucket> or mc mirror --overwrite --remove old/<bucket> new/<bucket>: run from any machine that reaches both endpoints, suited to the final pass while IOMETE is stopped.

An S3-level copy does not carry bucket policies, versioning or lifecycle settings, users or access keys. Set these up on the new storage before the cutover. It also copies only the current version of each object, which is all IOMETE needs.

Backend Copy With rsync​

Some teams copy the storage system's data directories at file level, for example with rsync. This only works when the new storage has the same layout as the old one: the same number of servers and drives, each drive copied to its matching drive, and the same storage software version. S3-compatible products such as MinIO spread each object, and their own users, policies and bucket settings, across several drives, so a file-level copy between different layouts produces a broken store. The storage service must also be stopped during the final copy.

Cutover Procedure​

At the moment of switching, the new storage must contain exactly the objects of the old storage, and the catalog in PostgreSQL must match both.

1. Prepare​

  1. Set up the new storage: buckets with the same names, the same access keys and policies, and for a backend copy the same layout and version.
  2. Check that the new storage is reachable and its certificate trusted from everywhere listed under Certificates and Network Access. When you keep the old host name, the certificate must be valid for that name.
  3. Prepare the switch: the DNS change, or for an endpoint change the updated Helm values and the changes listed under Settings Entered Manually.
  4. Record the "before" measurements in Table Level for your most important tables.
  5. Agree on a downtime window and inform users.

2. Run the Initial Copy​

Copy the data while the platform is in normal use, with one or more passes. The copy is not consistent yet, which is expected. It shortens the final pass.

3. Freeze Writes​

Stop everything that writes to storage or to the catalog:

  1. Suspend all scheduled Spark jobs through the console or API, not in Kubernetes, and keep a list of what you suspend.

  2. Stop streaming jobs and other long-running jobs.

  3. Terminate compute clusters, Spark Connect clusters and Jupyter containers.

  4. Disable automated table maintenance if it is enabled.

  5. Stop external tools that write to the storage or use the IOMETE catalog.

  6. Scale down the IOMETE services, keeping a record of their replica counts:

    kubectl get deployment,statefulset -n <namespace> -l app.kubernetes.io/instance=<release-name> > replicas-before-migration.txt
    kubectl scale deployment,statefulset -n <namespace> -l app.kubernetes.io/instance=<release-name> --replicas=0

    The label selects only the resources of the IOMETE release, so a PostgreSQL instance in the same namespace keeps running for the backup in the next step.

  7. Back up the IOMETE PostgreSQL databases. This backup matches the storage at this moment and is your rollback point.

4. Run the Final Copy​

  1. For a backend copy, stop the S3 service on the old storage first. A running storage service can leave its files inconsistent even when no one is writing.
  2. Run the final pass. Make sure it also deletes objects on the new storage that were removed from the old storage since the earlier passes (rclone sync, mc mirror --remove, rsync --delete).
  3. Start the S3 service on the new storage if it is stopped.

5. Switch​

  1. Switch to the new storage:
    • Keeping the host name: update the DNS record, then scale the IOMETE services back up to the replica counts in replicas-before-migration.txt.
    • Changing the endpoint: run the Helm upgrade with the new endpoint and any certificate changes. Services scaled down in the freeze start again with the new settings. Restore any that stay at zero replicas from replicas-before-migration.txt. Then apply the changes from Settings Entered Manually.
  2. Restart services and workloads.
  3. Run the validation checks.

6. Reopen​

  1. Resume the scheduled jobs you suspended, through the console or API. This also refreshes legacy jobs.
  2. For an endpoint change, refresh any legacy scheduled job that was not suspended, then confirm the legacy job check returns nothing.
  3. Restart streaming jobs, re-enable table maintenance and let external tools connect.

Rolling Back​

Keep the old storage unchanged, with its S3 service stopped or blocked from the cluster, and start it only for the object comparison or to roll back. A setting that still points to the old endpoint then fails with a clear connection error instead of silently reading out-of-date data. To roll back, point the DNS record back or restore the previous Helm values and run the Helm upgrade, then start the old storage and restart services and workloads.

warning

Rollback is simple only until workloads resume. After that, new data exists only in the new storage and a rollback loses it. Decide go or no-go before you reopen.

Validating the Migration​

Object Level​

Compare both storages through the S3 API, with the old storage started temporarily (read-only if possible) while no IOMETE workloads run.

rclone check --size-only old:<bucket> new:<bucket>
# or
mc diff old/<bucket> new/<bucket>

Object names and sizes must match. A backend copy also keeps each object's ETag, so you can drop --size-only to compare checksums. After an S3-level copy, ETags of objects uploaded in parts can differ even when the content is identical.

Catalog Level​

Every Iceberg table's current metadata file must exist on the new storage. This script reads the pointers from the Iceberg catalog database (iomete_iceberg_db with the default database prefix) and reports any that are missing. Run it once for every bucket you moved:

BUCKET="<bucket>"

psql -h <db-host> -U <db-user> -d iomete_iceberg_db -At \
-c "SELECT metadata_location FROM iceberg_tables" \
| grep -E "^s3a?://${BUCKET}/" \
| sed -E 's#^s3a?://##' \
| while read -r path; do
aws s3api head-object --endpoint-url "https://<new-endpoint-host>" \
--bucket "${path%%/*}" --key "${path#*/}" > /dev/null 2>&1 \
|| echo "MISSING: $path"
done

The expected output is empty.

Table Level​

Before the freeze, record these results for your important tables, then repeat them after the switch and compare:

-- Latest snapshot
SELECT snapshot_id, committed_at
FROM <catalog>.<database>.<table>.snapshots
ORDER BY committed_at DESC
LIMIT 1;

-- Data files, rows and bytes in the current version
SELECT count(*) AS data_files,
sum(record_count) AS total_rows,
sum(file_size_in_bytes) AS total_bytes
FROM <catalog>.<database>.<table>.files;

These queries read only metadata. To confirm the data files are readable, run a full read on your most critical tables. The result must be identical before and after:

SELECT count(*) AS row_count, sum(hash(*)) AS content_checksum
FROM <catalog>.<database>.<table>;

Platform Level​

CheckExpected result
Create a table, insert rows, read them back, drop the tableAll steps succeed
Query an existing table from the SQL EditorResults returned
Open a few saved SQL Editor worksheetsContent shown
For an endpoint change, run the checks under Finding Remaining ReferencesNo output
Let one scheduled Spark job run on its scheduleJob succeeds
Open the Spark History ServerApplications from before the migration are listed
Query from an external client (JDBC, Arrow Flight or a BI tool)Full result received
Read a table from an external engine that uses the IOMETE REST catalog, if you have oneData returned
Query a table protected by data access policiesAccess enforced and audit entry recorded