Isometric illustration of an old and a new app server both connected to the same database during a migration

Zero-downtime Rails migrations come down to two rules: never take a lock that blocks traffic for more than a moment, and never ship a schema change that the currently running code can’t handle. Zero-downtime deploys make the second rule unavoidable, because for a few seconds or minutes the old and new release run side by side against the same database. This playbook shows how to set lock timeouts, split risky changes into expand-and-contract steps, add indexes and constraints without blocking writes, backfill data safely, and let strong_migrations catch the mistakes before they reach production.


TL;DR

  • Set lock_timeout (5–10 s) for migrations. A migration waiting for a lock blocks every query queued behind it, so it should fail fast and be retried instead.
  • Use expand → migrate → contract: add new columns first, deploy code that handles both shapes, backfill, then remove the old structure in a later deploy.
  • Before removing a column, add it to ignored_columns and deploy. Otherwise running processes raise PG::UndefinedColumn on their next insert.
  • On PostgreSQL, add indexes with algorithm: :concurrently and disable_ddl_transaction!. Add foreign keys and check constraints with validate: false, then validate them in a separate migration.
  • Backfill in batches from a background job, never inside the migration transaction, and install strong_migrations to block unsafe operations in development.

Table of Contents

Why do Rails migrations cause downtime?

Rails migrations cause downtime in two ways: a DDL statement takes a lock that blocks reads or writes on a busy table, or the schema changes in a way the still-running old code doesn’t expect. Both are avoidable.

Locks. Most ALTER TABLE statements in PostgreSQL need an ACCESS EXCLUSIVE lock, which conflicts with everything, including plain SELECTs. The ALTER itself may take milliseconds. The danger is waiting for the lock:

Why a fast migration can still take the site down: lock queue behind a long-running query

If a long-running query holds even a weak lock on the table, the migration waits. While it waits, every new query on that table queues behind the migration. Puma threads fill up waiting on PostgreSQL, the health check times out, and users see 503s, all because of a migration that would have taken 4 ms.

Code and schema overlap. During a zero-downtime deploy, the new release boots while the old one is still serving traffic. If the migration removes or renames a column, the old processes keep using the old column name and fail. Active Record also caches each table’s columns when a model first loads, so removing a column breaks INSERTs from processes that started before the migration, even if no code reads that column.

Prerequisites

  • Rails 7.1+ (examples use Rails 8.0 and Ruby 3.3).
  • PostgreSQL 12 or later. Some safe patterns below depend on PostgreSQL 11 and 12 behaviour; notes indicate which.
  • A deploy process where migrations run before or during the switch to the new release (Step 6).
  • A job backend such as Solid Queue or Sidekiq for backfills.

Step 1: Install strong_migrations and set timeouts

strong_migrations checks each migration against a list of unsafe operations and refuses to run it with an explanation and a safe alternative. Add it before you need it.

bundle add strong_migrations
bin/rails generate strong_migrations:install

Edit the generated initializer:

# config/initializers/strong_migrations.rb
# Give up quickly if a lock isn't available instead of queueing traffic behind us.
StrongMigrations.lock_timeout = 10.seconds

# Long enough for real DDL, short enough to catch runaway statements.
StrongMigrations.statement_timeout = 1.hour

# Retry lock timeouts automatically (only for migrations that are safe to retry).
StrongMigrations.lock_timeout_retries = 3

# Match your production server so version-specific checks are correct.
StrongMigrations.target_version = 16

# Only check migrations created after the gem was installed.
StrongMigrations.start_after = 20261011000000

Also set timeouts for the application’s own connections, so a slow report can’t hold a lock long enough to cause the pile-up described above:

# config/database.yml
production:
  primary:
    url: <%= ENV["DATABASE_URL"] %>
    variables:
      statement_timeout: 15s
      lock_timeout: 5s

variables are sent as SET commands on every new connection. Migrations override them because strong_migrations sets its own values at the start of each run.

If you don’t want the gem, set the lock timeout by hand at the top of each risky migration:

# db/migrate/20261011090000_add_status_to_orders.rb
class AddStatusToOrders < ActiveRecord::Migration[8.0]
  def change
    execute "SET lock_timeout = '5s'"
    add_column :orders, :status, :string
  end
end

Step 2: Add columns safely

Adding a nullable column, or a column with a constant default, is safe on PostgreSQL 11+ because it only changes the catalog. It doesn’t rewrite the table. The lock is held for milliseconds, and the lock_timeout from Step 1 protects you from the queueing problem.

# db/migrate/20261011090100_add_status_to_orders.rb
class AddStatusToOrders < ActiveRecord::Migration[8.0]
  def change
    add_column :orders, :status, :string, default: "pending"
  end
end

What still rewrites the table, and must be avoided on large tables:

Operation Safe on PG 11+? Why
add_column nullable, no default Yes Catalog-only change
add_column with a constant default Yes (PG 11+) Default stored in the catalog
add_column with a volatile default (gen_random_uuid(), now()) No Every row needs its own value: full rewrite
add_column ... null: false without default No on non-empty tables Fails or rewrites; use the NOT NULL pattern in Step 5
change_column type (e.g. integer → bigint) No Full rewrite under ACCESS EXCLUSIVE
change_column varchar(50) → varchar(100) or → text Yes Widening without rewrite

For a volatile default, add the column without it, set the default for new rows with change_column_default, and backfill existing rows (Step 4).

Pro Tip: Keep each migration to one DDL change on a large table. If the third statement times out on a lock, you don’t want the first two rolled back and re-run in the middle of a deploy.

Step 3: Remove and rename columns with expand and contract

Never remove or rename a column in the same deploy as the code change. Tell Active Record to ignore the column first, deploy, and only then drop it.

Removing a column

Deploy 1: stop using the column and ignore it:

# app/models/user.rb
class User < ApplicationRecord
  self.ignored_columns += ["legacy_role"]
end

Deploy 2: drop it. strong_migrations requires safety_assured here because it can’t know the column is already ignored:

# db/migrate/20261012080000_remove_legacy_role_from_users.rb
class RemoveLegacyRoleFromUsers < ActiveRecord::Migration[8.0]
  def change
    safety_assured { remove_column :users, :legacy_role, :string }
  end
end

Deploy 3 (any later deploy): remove the ignored_columns line.

Renaming a column

rename_column is instant in PostgreSQL, but it breaks every running process that uses the old name. Use the expand-and-contract sequence instead:

Expand, migrate, contract: renaming a column over three deploys

Deploy 1 (expand): add the new column and write to both.

# db/migrate/20261011091000_add_full_name_to_users.rb
class AddFullNameToUsers < ActiveRecord::Migration[8.0]
  def change
    add_column :users, :full_name, :string
  end
end
# app/models/user.rb
class User < ApplicationRecord
  before_save :sync_full_name

  private

  def sync_full_name
    self.full_name = name if will_save_change_to_name?
  end
end

Backfill existing rows (Step 4).

Deploy 2: read from full_name everywhere, keep writing both, and ignore the old column.

Deploy 3 (contract): remove_column :users, :name inside safety_assured, and delete the sync callback.

Renaming a table follows the same pattern, with a new table, dual writes and a backfill. If the table is small and the rename is worth the trouble, a PostgreSQL view with the old name can bridge the gap instead.

Step 4: Backfill data in batches

Backfill from a background job in small batches, outside of any migration transaction. A single UPDATE users SET full_name = name on ten million rows holds row locks for the whole statement, generates huge WAL traffic and can run into statement_timeout.

# app/jobs/backfill_full_name_job.rb
class BackfillFullNameJob < ApplicationJob
  queue_as :low

  BATCH_SIZE = 1_000

  def perform
    User.where(full_name: nil).in_batches(of: BATCH_SIZE) do |batch|
      batch.update_all("full_name = name")
      sleep 0.05 # let replicas and autovacuum keep up
    end
  end
end

Enqueue it once the expand deploy is live:

bin/rails runner 'BackfillFullNameJob.perform_later'

This job is safe to rerun: where(full_name: nil) skips rows that are already done, so a deploy or crash that interrupts it just means you enqueue it again. That’s the at-least-once rule from our background jobs guide.

If you’d rather keep backfills in migrations, for example so they run in every environment, disable the transaction and batch inside the migration:

# db/migrate/20261011093000_backfill_users_full_name.rb
class BackfillUsersFullName < ActiveRecord::Migration[8.0]
  disable_ddl_transaction!

  class MigrationUser < ActiveRecord::Base
    self.table_name = "users"
  end

  def up
    MigrationUser.unscoped.where(full_name: nil).in_batches(of: 1_000) do |batch|
      batch.update_all("full_name = name")
      sleep 0.05
    end
  end

  def down
    # no-op: data backfill
  end
end

The inline model class means the migration doesn’t depend on app/models/user.rb, which may have changed or disappeared by the time someone runs it.

Step 5: Add indexes, foreign keys and NOT NULL without blocking

Each of these has a blocking default and a non-blocking alternative in PostgreSQL. The non-blocking version always takes two steps: create the object without checking existing rows, then validate it separately under a weaker lock.

Indexes

# db/migrate/20261011094000_add_index_on_orders_status.rb
class AddIndexOnOrdersStatus < ActiveRecord::Migration[8.0]
  disable_ddl_transaction!

  def change
    add_index :orders, :status, algorithm: :concurrently
  end
end

CREATE INDEX CONCURRENTLY doesn’t block writes, but it can’t run inside a transaction, which is why you need disable_ddl_transaction!. The same applies to new references: add_reference :orders, :coupon, index: { algorithm: :concurrently }.

Foreign keys

# db/migrate/20261011095000_add_foreign_key_orders_users.rb
class AddForeignKeyOrdersUsers < ActiveRecord::Migration[8.0]
  def change
    add_foreign_key :orders, :users, validate: false
  end
end
# db/migrate/20261011095100_validate_foreign_key_orders_users.rb
class ValidateForeignKeyOrdersUsers < ActiveRecord::Migration[8.0]
  def change
    validate_foreign_key :orders, :users
  end
end

The first migration enforces the constraint for new rows straight away. The second scans existing rows under a SHARE UPDATE EXCLUSIVE lock, so reads and writes keep working.

Making a column NOT NULL

change_column_null :users, :full_name, false scans the whole table under ACCESS EXCLUSIVE. On PostgreSQL 12+, a validated check constraint lets SET NOT NULL skip that scan:

# db/migrate/20261012090000_add_full_name_not_null_check.rb
class AddFullNameNotNullCheck < ActiveRecord::Migration[8.0]
  def change
    add_check_constraint :users, "full_name IS NOT NULL", name: "users_full_name_null", validate: false
  end
end
# db/migrate/20261012090100_validate_full_name_not_null.rb
class ValidateFullNameNotNull < ActiveRecord::Migration[8.0]
  def up
    validate_check_constraint :users, name: "users_full_name_null"
    change_column_null :users, :full_name, false
    remove_check_constraint :users, name: "users_full_name_null"
  end

  def down
    add_check_constraint :users, "full_name IS NOT NULL", name: "users_full_name_null", validate: false
    change_column_null :users, :full_name, true
  end
end

Run the backfill before the second migration, or validation fails on the rows that are still NULL.

Change Blocking default Non-blocking pattern
Add index add_index blocks writes algorithm: :concurrently + disable_ddl_transaction!
Add foreign key Locks both tables while it scans validate: false, then validate_foreign_key
Add check constraint Blocks reads/writes while it scans validate: false, then validate_check_constraint
Set NOT NULL Full scan under ACCESS EXCLUSIVE Validated check constraint first (PG 12+)
Change column type Full rewrite New column + backfill + swap (expand/contract)

Step 6: Decide where migrations run during a deploy

Run migrations once per release, before the new code takes traffic, while the old release keeps serving. With the expand-and-contract discipline above, the old code is compatible with the new schema, so this ordering is safe.

The Rails 8 Docker entrypoint already runs db:prepare when a container starts the web server, before it answers the health check:

#!/bin/bash -e
# bin/docker-entrypoint (excerpt)
if [ "${@: -2:1}" == "./bin/rails" ] && [ "${@: -1:1}" == "server" ]; then
  ./bin/rails db:prepare
fi
exec "${@}"

That works well with Kamal and other container deploys: the new container migrates, then boots Puma, then passes /up, and only then does the proxy switch traffic. Rails takes a PostgreSQL advisory lock while migrating, so two containers starting together won’t run the same migration twice. The full entrypoint is explained in how to dockerize a Rails app, and the health check side in Rails health check endpoints.

Make sure your deploy timeout covers the migration, though. Kamal waits deploy_timeout (30 s by default) for the new container to turn healthy. A 2-minute index build in the entrypoint fails the deploy even though the migration itself succeeds. For long migrations, run them as a separate step before deploying:

kamal build push
kamal app exec --primary --version=$(git rev-parse HEAD) "bin/rails db:migrate"
kamal deploy --skip-push

Step 7: Verify it works

Check that each migration is caught or allowed by strong_migrations in development, then watch locks in production while it runs.

bin/rails db:migrate

An unsafe migration fails like this before touching the database:

=== Dangerous operation detected #strong_migrations ===

Active Record caches attributes, which causes problems
when removing columns. Be sure to ignore the column:

class User < ApplicationRecord
  self.ignored_columns += ["legacy_role"]
end

During the production run, watch for sessions waiting on locks from a second terminal:

-- psql "$DATABASE_URL"
SELECT pid,
       now() - query_start AS waiting_for,
       wait_event_type,
       left(query, 70) AS query
FROM pg_stat_activity
WHERE wait_event_type = 'Lock'
ORDER BY query_start;

An empty result, or rows that clear in under a second, means the migration isn’t blocking traffic. After a concurrent index build, confirm the index is valid:

SELECT indexrelid::regclass AS index, indisvalid
FROM pg_index
WHERE indrelid = 'orders'::regclass;

How Heroku, Render, Fly.io, Upsun and Enkihost run migrations

Every platform provides a hook that runs migrations once, after the build and before the new release takes traffic. None of them makes the migration itself safe. That’s still your job, using the patterns above.

Platform Hook Runs Timeout On failure
Heroku release: in Procfile (release: bin/rails db:migrate) One-off dyno before new dynos boot 1 hour, not extendable Release not deployed; old dynos keep running
Render Pre-deploy command Separate instance after build, before deploy (paid services only) 30 minutes Deploy fails; previous deploy keeps serving
Fly.io release_command in [deploy] of fly.toml Temporary Machine from the new image Configurable Deploy aborted
Upsun hooks.deploy On the app container during deploy No documented hard limit Idempotent requests (GET, PUT, DELETE) are held during the hook; POST and PATCH are not
Kamal 2 Docker entrypoint db:prepare, or kamal app exec New container before it turns healthy deploy_timeout (30 s default) New container never healthy; old one keeps serving
Enkihost No release phase; db:prepare in bin/docker-entrypoint New container, before Puma starts Not documented Container exits, deploy aborted; old container keeps serving

The Upsun detail is worth calling out: because requests are held while the deploy hook runs, a slow migration there shows up as slow page loads rather than errors, and non-idempotent requests aren’t held at all. Short, lock-safe migrations matter as much there as anywhere else.

# fly.toml
[deploy]
  release_command = "./bin/rails db:migrate"
# Procfile (Heroku)
release: bin/rails db:migrate
web: bundle exec puma -C config/puma.rb

Enkihost does zero-downtime deploys, which means old and new code overlap exactly as described in this guide. The expand-and-contract rules apply there unchanged.

Troubleshooting migration failures

These are the errors you’ll see when a migration hits a lock, runs inside a transaction it can’t use, or races with running code.

PG::LockNotAvailable: canceling statement due to lock timeout

ActiveRecord::LockWaitTimeout: PG::LockNotAvailable: ERROR:  canceling statement due to lock timeout

This is lock_timeout doing its job: something held a lock on the table for longer than the timeout. Find the blocker with the pg_stat_activity query above (often a long report, a stuck transaction or an idle in transaction session), then retry. StrongMigrations.lock_timeout_retries does the retry for you.

CREATE INDEX CONCURRENTLY cannot run inside a transaction block

PG::ActiveSqlTransaction: ERROR:  CREATE INDEX CONCURRENTLY cannot run inside a transaction block

Add disable_ddl_transaction! to the migration class. Keep that migration to the index only, since nothing else in it is wrapped in a transaction either.

PG::UndefinedColumn right after a deploy

ActiveRecord::StatementInvalid: PG::UndefinedColumn: ERROR:  column "legacy_role" of relation "users" does not exist

A process that started before the migration still has the old column list cached and includes it in INSERTs. Use ignored_columns and deploy it before the migration that drops the column.

ActiveRecord::ConcurrentMigrationError

ActiveRecord::ConcurrentMigrationError: Cannot run migrations because another migration process is currently running.

Two processes started migrating at the same time, such as two containers booting together or a release phase plus an entrypoint. The advisory lock protected you. Pick one place to run migrations, and make the other path skip them.

An invalid index left behind

If a concurrent index build fails (a lock timeout, or a duplicate in a unique index), PostgreSQL leaves an INVALID index that slows down writes and isn’t used by reads. Drop it and try again:

# db/migrate/20261011094500_retry_index_on_orders_status.rb
class RetryIndexOnOrdersStatus < ActiveRecord::Migration[8.0]
  disable_ddl_transaction!

  def up
    remove_index :orders, :status, algorithm: :concurrently, if_exists: true
    add_index :orders, :status, algorithm: :concurrently
  end

  def down
    remove_index :orders, :status, algorithm: :concurrently, if_exists: true
  end
end

PG::QueryCanceled: canceling statement due to statement timeout in a backfill

The batch is too large or the WHERE isn’t indexed. Lower the batch size, add an index on the column you filter by, or batch by primary key ranges.

Author Perspective: boring migrations, many deploys

The biggest change in how I handle migrations wasn’t a tool. It was accepting that a schema change can take three deploys. It feels slow the first time. Then you notice that each deploy is small, each one can be rolled back, and none of them needs a maintenance window or a 2 AM slot. strong_migrations is the first gem I add to every Rails app that has real users, because I’d rather argue with a linter in development than with a lock queue in production. If a migration needs safety_assured, I write a comment explaining why. Future me always asks.

Running migrations on Enkihost

Zero-downtime migrations depend on two things: a deploy process where old and new code overlap without dropping requests, and a PostgreSQL database you can reason about. On Enkihost, the patterns in this guide work as written:

Enkihost

  • Zero-downtime deploys: the new release takes over without dropping in-flight requests, so expand-and-contract migrations keep both releases working during the overlap.
  • PostgreSQL add-on injects DATABASE_URL, so the database.yml variables for statement_timeout and lock_timeout apply to every connection your app opens.
  • Per-app resource isolation: your backfill jobs use your app’s allocated CPU and memory, and nobody else’s traffic slows them down.

Rails and Sinatra apps run on Ignite (5 EUR per month after a 14-day free trial, with PostgreSQL and Redis included) or Blaze (16 EUR per month, with high availability and autoscaling). The free Spark plan is for Jekyll sites. See enkihost.com.

FAQ

Is adding a column with a default value safe in PostgreSQL?

Yes, on PostgreSQL 11 and later, as long as the default is a constant. The default is stored in the catalog and the table isn’t rewritten. A volatile default such as now() or gen_random_uuid() still rewrites every row, so add those columns without a default and backfill.

Should I run migrations before or after deploying the new code?

Run them before the new code takes traffic, while the old code is still serving. That only works if every migration is backwards compatible with the old code, which is what the expand and contract pattern guarantees. Destructive changes such as dropping a column go in a later deploy.

Can I roll back a zero-downtime migration?

Usually yes, because expand-style migrations only add things. Rolling back the code is safe since the old code ignores the new column. Contract migrations that drop data are not reversible in practice, which is why they come last and only after the new code has run in production for a while.

Does strong_migrations work with MySQL?

Yes. strong_migrations supports PostgreSQL, MySQL and MariaDB, and adjusts its checks to the database and version you set in target_version. Some safe alternatives differ; for example, MySQL uses online DDL algorithms instead of concurrent index builds.

How long should lock_timeout be for migrations?

Between 5 and 10 seconds is a good default for web apps. It must be shorter than the time your users and health checks will tolerate requests queueing, and long enough that the migration gets its lock on a normally busy table. Combine it with automatic retries.

Sources