I’ve been rebuilding the deployment process for an application recently, and the original idea was fairly straightforward. Each version gets built into its own immutable release directory, everything is prepared away from the live application, and once it’s ready you switch a /current symlink over to the new release.

I’ve used variations of that setup plenty of times before and I really like it. You get a nice clean boundary between releases, you’re not modifying the application that’s currently serving traffic, and you avoid that particularly exciting deployment strategy where production spends a few seconds being made up of half the old release and half the new one.

Something along these lines:

/var/www/app/
├── releases/
│ ├── a1b2c3/
│ ├── d4e5f6/
│ └── g7h8i9/
├── shared/
└── current -> releases/d4e5f6/

You build the next release somewhere else, install its dependencies, render its configuration, build the caches and compiled views, validate it, and then move /current when you’re happy.

The symlink bit is easy.

Then I got to the database.

This particular application has a central database as well as a separate database for each tenant. It also has Horizon workers, scheduled jobs and other processes that can carry on running long after the web request that originally started them has disappeared. Once I actually started working through how all of those pieces behave during a release, describing the deployment as “atomic” started feeling a bit dishonest.

The code switch can be atomic. The application around it can’t.

The filesystem isn’t really the difficult bit

One of the nice things about immutable release directories is that preparing the new application doesn’t need to affect the one that’s currently live. Release B can have its own dependencies, caches, compiled views and generated configuration while release A carries on doing its thing.

I actually tightened that isolation while working on this because there were still a few bits of generated application state being shared between releases. That sort of sharing looks convenient until preparing B mutates something that A is still relying on, at which point you’ve quietly undermined quite a lot of the reason for using immutable releases in the first place.

Once everything is self-contained, activating the application itself becomes almost boring. /current points at release A, you change it to release B, restart whatever processes need to load the new code and you’re done.

Code rollback is similarly nice. Point /current back at A and restart the relevant processes.

Unfortunately, the database isn’t sitting inside releases/d4e5f6/, so it doesn’t get to participate in any of that.

There is always a point where the versions overlap

Say release A is currently running and release B includes a database migration. A fairly sensible deployment sequence looks like this:

prepare release B
↓
run migrations
↓
switch /current to B

The problem is that there is necessarily some amount of time between the migration completing and the application switching over. During that period, release A is still running against a database that has already been changed for release B.

Sometimes that’s absolutely fine. If the migration just adds a nullable column that release A knows nothing about, who cares? The old application ignores it and carries on.

But if release B renames a column, removes something release A still uses, changes a constraint or otherwise makes the schema incompatible with the currently running code, you’ve created a state that was never really intended to exist:

release A code
+
release B database

The symlink can still have changed atomically. It doesn’t really help.

That was the point where I stopped thinking about the deployment as “build some files and swap a symlink” and started looking at it as a transition between two versions of a whole system. For the transition to be safe, there has to be some amount of compatibility between the state before the switch and the state afterwards.

One database would have made this considerably easier

The tenant databases make the problem a lot more obvious.

A release with schema changes doesn’t run one migration. It needs to migrate the central database and then every active tenant database. If there are a hundred tenants, that means there are potentially a hundred and one separate points where something can fail.

Suppose the central database succeeds, tenant databases 1 through 36 succeed, and tenant 37 fails.

You now have part of the platform on the new schema, part of it on the old schema, and the old application is still the active release because the deployment hasn’t reached activation yet.

There’s no clever bit of automation that suddenly makes that state transactional. It’s partial by definition.

One of the things I ended up tightening while rebuilding this was how failures from those tenant migrations propagate. If one of the nested migration commands fails, that failure has to make it all the way back to the process controlling the deployment.

This sounds painfully obvious, but it’s surprisingly easy to end up with something like a shell loop where a command prints an error for tenant 37, the loop carries on doing other work, and the outer script eventually exits successfully. That’s irritating when you’re running it manually. Once you’ve got automation trusting that exit code, it’s a genuine deployment bug.

For this application, the rule became pretty simple: if one active tenant migration fails, the migration phase has failed and the new release does not get activated.

That doesn’t magically undo the migrations that have already succeeded, but at least the deployment isn’t pretending partial success is good enough and moving production onto the new code anyway.

Maintenance mode doesn’t stop your application

Another assumption I had to get rid of was that putting Laravel into maintenance mode somehow means the application has stopped.

It doesn’t.

It stops normal HTTP traffic reaching the application, which is useful, but HTTP requests are only one of the things writing to the database. Horizon workers are still processing jobs. The scheduler can still start work. There may be other long-running processes around as well, and all of those can be running code that was loaded before the deployment started.

The symlink changing underneath a PHP process does not magically replace the code already loaded into that process either. A worker started from release A can quite happily carry on executing release A code while /current now points at release B.

So before crossing the migration boundary, I needed an actual quiescence phase rather than just maintenance mode. The scheduler, Horizon and PHP-FPM are stopped in a controlled order, the application is put into maintenance, and only then do the migrations begin. Once the migrations have completed and the new release is active, those processes can be started again against the new version.

I added a deployment lock around the whole operation as well, because two completely valid deployments running at the same time can produce a spectacularly invalid result if both believe they own the migration and activation sequence.

The thing that matters isn’t stopping web traffic. It’s stopping every old process that can still write using the old application’s assumptions.

That’s a much more useful boundary.

I spent more time thinking about failure than success

The happy path for a deployment is easy to write down:

build
migrate
activate
restart
health check

Every command succeeds, the site comes back, everyone goes home. Lovely.

The more useful exercise was taking each step and asking “okay, if this one fails, what state are we actually in now?”

If building or validating release B fails, that’s easy. Nothing live has changed yet, so the deployment can stop and production carries on running release A.

A migration failure is more awkward. The symlink is still pointing at release A, but I can’t necessarily say that the database is still in release A’s state because some of the migrations may already have completed. Automatically taking the application back out of maintenance and restarting all of the old writers would therefore be making a fairly big assumption: that release A is safe against whatever database state we’ve just failed to finish creating.

I don’t think the deployment should make that assumption.

So in that situation it stops. Maintenance stays enabled, the writers remain stopped and somebody has to look at what actually happened.

That’s obviously less convenient than a deploy script heroically trying to recover everything itself, but I would much rather have production stopped in a state I can understand than running in a state the automation has guessed is probably okay.

If the application has already activated and the health check then fails, moving /current back to the previous release may be possible. But again, whether that is actually a valid rollback depends entirely on what happened to the database.

The deployment now checks the release link after activation as well rather than simply assuming that because the symlink command returned successfully, reality definitely matches the plan.

None of those checks are particularly clever in isolation. The important bit is that the deployment has stopped treating “the last command exited zero” as proof that the whole system is in the state we think it is.

A failed deployment should leave me somewhere I can explain.

Rollback isn’t the same thing as going backwards

Immutable releases make application rollback look wonderfully simple.

current -> release B

becomes:

current -> release A

And for the application files, that’s genuinely great.

But if release B changed the database in a way release A doesn’t understand, you haven’t restored the previous system. You’ve restored the previous code and left the database in the future.

I definitely don’t want a failed health check automatically running every available down() migration and hoping that puts everything back exactly how it was either. Some migrations are destructive, some transform data, and some technically have a reverse operation while still being something I would never want happening automatically on a production database.

Depending on what failed, the correct recovery might be moving forward with a fix. It might be a deliberate restore from a database backup. It might be reverting some specific schema change. There isn’t one generic “undo deployment” operation that safely handles all of those cases.

This has made me value compatibility between adjacent releases a lot more than clever rollback automation.

For bigger schema changes, the boring expand-and-contract approach buys you a lot. Add the new structure first while leaving the old one in place. Deploy application code that understands both. Move everything over to the new structure, then remove the old one in a later release once nothing relies on it anymore.

Yes, that means what could theoretically have been one migration becomes two or three deployments. But it also means release A and release B can actually coexist for a while, which turns out to be quite useful when your deployment process literally requires them to coexist for a while.

The web application probably shouldn’t be able to do migrations anyway

While rebuilding all of this I ended up looking at the database permissions as well, because the deployment work exposed another thing I wasn’t particularly happy with.

The normal application process needs to read and modify application data. That doesn’t mean it needs enough MySQL privileges to perform every schema change and tenant lifecycle operation the platform might ever need.

Those are now separate paths. The runtime database identity has the permissions the application needs day to day, while migrations and tenant lifecycle operations use a separate privileged identity with the extra authority needed for those jobs.

It doesn’t solve the deployment problem by itself, but it follows the same thinking as most of the other changes: if something exceptional needs more authority, make that an explicit boundary rather than just giving the normal process everything because it’s convenient.

If somebody compromises the web application, the fact they can modify application data is already bad enough. I’d rather not also hand them the ability to redesign the database while they’re there.

Ansible didn’t actually solve any of this

A lot of the deployment logic started out spread across shell scripts, and part of this work has been moving the infrastructure towards Ansible as the authoritative orchestration layer.

I think that’s the right choice because the deployment has gradually become much more about enforcing state than simply running a list of commands. I want a release to exist in a particular form, configuration to belong to that release, particular processes to be stopped before migrations begin, every required migration to complete before activation, and the correct release to actually be current afterwards.

Ansible is a much nicer fit for that than allowing a collection of shell scripts to slowly evolve into its own slightly haunted configuration-management system.

But Ansible isn’t really what solved the interesting part.

I could have implemented the same deployment in shell, Python, Go or pretty much anything else. The difficult bit was deciding what the valid states of the system actually are, which transitions between them are safe, and where the deployment needs to stop instead of trying to be clever when reality doesn’t match what I expected.

Once you’ve worked that out, the orchestration tool mostly just makes sure those rules get followed.

I want activation to be the boring bit

The deployment now does quite a lot before /current ever moves, which is exactly how I want it.

The next release gets built away from production and owns its generated state. It gets validated before anything live changes. A deployment lock makes sure another release isn’t trying to cross the same boundary at the same time. Existing writers are stopped before the database changes, and both the central migration and every required tenant migration have to succeed before activation is allowed to happen.

Only then does the symlink move.

After that, the processes come back against the new release, the link is checked to make sure the version I think is live actually is live, and health checks determine whether the application works rather than whether a directory happens to exist.

The symlink switch is still atomic, and that’s useful. I just don’t think it’s particularly helpful to describe the entire deployment that way. There is no single transaction wrapping application files, a central database, potentially hundreds of tenant databases, PHP-FPM, Horizon, the scheduler and everything else that’s involved in keeping the application alive.

Those intermediate states exist whether I acknowledge them or not.

What I actually want from the deployment is much less magical. I want adjacent releases to be compatible wherever that’s practical, I want to know exactly which old processes are capable of touching the new state, and I want the deployment to stop when it reaches a situation it can’t prove is safe rather than carrying on because that’s what the happy-path script says comes next.

By the time /current moves, all of the difficult deployment work should already be finished.

The switch itself should be boring.