
Geschreven door Funs Janssen
Software Consultant
Ik ben Funs Janssen. Ik ontwikkel software en schrijf over de beslissingen die daarbij komen kijken: architectuur, ontwikkelmethoden, AI-tools en de zakelijke impact van technische keuzes. Deze blog is een verzameling praktische aantekeningen van echte projecten: wat schaalbaar is, wat misgaat en wat in blogvriendelijke voorbeelden vaak over het hoofd wordt gezien.
Your pipeline was green yesterday. Today the integration tests fail with Could not find a valid Docker environment or Cannot connect to the Docker daemon at unix:///var/run/docker.sock. Nobody touched the YAML. Nobody touched the Dockerfile. And the same pipeline on your self-hosted pool still passes.
Your pipeline didn't change. The agent did. Starting with Azure Pipelines agent 5.279.0, Linux container jobs no longer get the host Docker socket mounted by default. Any job that runs Docker from inside a container job, including every Testcontainers suite, loses its connection to the daemon.
In this post you get the exact fix, a short guide to deciding whether you should turn the socket back on at all, the temporary escape hatch for self-hosted agents, and a way to find every affected pipeline before the rest of your agents update.
What changed in agent 5.279.0
From "always mapped" to opt-in
For years, the Azure Pipelines agent quietly mounted /var/run/docker.sock into every Linux container job. You never asked for it, but it was there, so docker build, docker compose and Testcontainers all worked inside the job container.
The Sprint 279 release notes changed that: from agent version 5.279.0, the socket is not mapped by default. You now opt in per container with a new property, mapDockerSocket.
Why Microsoft did it
The Docker socket is not a harmless file. Whoever can talk to it can start a privileged container, mount the host filesystem and act as root on the machine. Microsoft's container jobs documentation says it plainly: code inside the container can run as root on your Docker host.
People had questioned the old default for years. A long-running issue on the agent repository asked why the socket was mounted at all, since most container jobs never use it. The change is the right call. It just arrived as one line in a sprint update most of us skim.
How to recognise it
The error messages
The symptom depends on what touched Docker first. The ones I see most:
- Docker CLI:
Cannot connect to the Docker daemon at unix:///var/run/docker.sock. Is the docker daemon running? - Testcontainers (Java, .NET, Node):
Could not find a valid Docker environmentor a timeout while looking for a Docker host - Docker Compose: a connection refused or "permission denied" style error on the socket path
- Build tasks that call
dockerfrom a script step inside the container job
If the job runs in a container: and the failing step talks to Docker, this change is your prime suspect.
Why hosted fails while self-hosted still passes
Microsoft-hosted agents pick up new agent versions first. Your self-hosted pool updates on its own schedule, or never if auto-update is off. That is why the same YAML can be red on ubuntu-latest and green on your own VMs. The Apache Flink project hit exactly this: fork builds on hosted agents broke while their self-hosted agents, still on an older version, kept passing until they upgraded.
To confirm, open the failing run and look at the Initialize job step. It logs the agent version. Anything at 5.279.0 or higher has the new default.
The fix: mapDockerSocket: true
On a container resource
If the job really needs Docker inside the container, opt in explicitly on the container resource:
1resources:2 containers:3 - container: build4 image: mcr.microsoft.com/dotnet/sdk:9.05 mapDockerSocket: true67jobs:8- job: IntegrationTests9 pool:10 vmImage: ubuntu-latest11 container: build12 steps:13 - script: dotnet test tests/Integration
That single property restores the old behaviour for this container only. The container resource schema documents it with a default of false.
Inline container jobs
If you declare the container inline on the job, the same property applies:
1jobs:2- job: IntegrationTests3 pool:4 vmImage: ubuntu-latest5 container:6 image: mcr.microsoft.com/dotnet/sdk:9.07 mapDockerSocket: true
One warning from my own reading: at the time of writing, the inline jobs.job.container schema page still describes mapDockerSocket as a flag you set to false to disable the mount, which reflects the old default. Trust the agent version and the container jobs page, not that description. If you rely on the socket, set true explicitly everywhere.
Templates: fix it once
Most organizations I work with define containers in a shared template repository. Fix it there, in the one place every pipeline pulls from, rather than in forty YAML files. A parameter works well:
1parameters:2- name: needsDocker3 type: boolean4 default: false56resources:7 containers:8 - container: build9 image: mcr.microsoft.com/dotnet/sdk:9.010 mapDockerSocket: ${{ parameters.needsDocker }}
Teams that need Docker pass needsDocker: true. Everyone else gets the safer default without doing anything.
Should you turn it back on? A quick decision guide
When the socket is justified
Mapping the socket is reasonable when:
- Testcontainers is your integration test strategy and the tests must run inside a specific SDK image
- You build or push images from inside a toolchain container
- You run docker compose to spin up a test environment as part of the job
In those cases, set mapDockerSocket: true on that container and move on. Make it a conscious, reviewed choice rather than an inherited default.
Safer alternatives
Before you flip the switch, check whether the job needs to be a container job at all:
- Run Docker steps on the host. A Microsoft-hosted Ubuntu agent already has Docker. If the only reason for the container was a specific SDK version, a
UseDotNet@2orNodeTool@0step on the host often does the same job, and Testcontainers works without any socket juggling. - Use service containers. If Testcontainers only starts Postgres or Redis, declare them as
services:on the job. The agent starts them next to your job container and you reach them by name, no socket needed. - Split the job. Build and test in the container, then build the image in a separate job on the host.
Self-hosted agents: know what you expose
On Microsoft-hosted agents the VM is thrown away after each job, so the blast radius of a mapped socket is small. On a self-hosted agent that runs many jobs on the same VM, the socket is a path from any pipeline's code to root on a machine that may hold cached credentials, other teams' workspaces and network access to internal systems. If you map the socket there, restrict which projects can use that pool.
Organization-wide: the temporary escape hatch
Restore the old default on agents you run
If you run a large self-hosted fleet and cannot update every YAML file this week, Microsoft provides a short-term switch. Set this environment variable on the agent host:
1AZP_AGENT_DEFAULT_MAP_DOCKER_SOCKET_TO_FALSE=false
With that set, Linux container jobs map the socket again unless a pipeline explicitly sets mapDockerSocket. This only works on agents you control; you cannot set host variables on Microsoft-hosted agents.
Give it an expiry date. Put a removal date in your agent provisioning scripts or a ticket on the platform backlog. A temporary security exception without an owner becomes permanent by accident.
Find every affected pipeline before your agents update
Search your repos
You are looking for two things in the same YAML: a container job, and a reason to talk to Docker. A rough but effective pass over a shallow clone of every repository in a project:
1az repos list --project MyProject --query "[].remoteUrl" -o tsv |2while read url; do git clone --depth 1 "$url"; done34grep -rlE "^\s*container:" --include="*.yml" --include="*.yaml" . |5xargs grep -liE "docker|testcontainers|compose"
That gives you a candidate list. Also check the test code itself: a Testcontainers dependency in a .csproj, pom.xml or package.json behind a container job is a hit even if the YAML never says "docker".
Control when self-hosted agents update
Agent pools have a setting that lets agents update themselves automatically. For a critical pool, consider turning that off temporarily, updating one canary agent first, and running your container pipelines on it. When those pass, let the rest of the pool update. I described the same "inventory first, then upgrade" approach for the Node 16 task end of life, and it works just as well here.
Stop agent changes from surprising you
This is not the last agent default that will change. Node runners are being retired, pipelines are moving to Microsoft Entra-issued tokens, and billing models are shifting, as I covered in pay-as-you-go agents vs parallel jobs. Three habits keep the surprises small:
- Read the sprint release notes once per sprint, filtered to Pipelines. It takes ten minutes.
- Keep a canary pipeline that exercises your container jobs, Testcontainers and image builds, and run it on a schedule against the newest agents.
- Put "agent and task changes reviewed" in your release checklist. Not as a vague intention but as an item someone ticks before a release goes out.
That last one is where I use my own tool. The free Checklist Extension for Azure DevOps lets you attach a reusable release-readiness checklist to a work item and block the state transition until it is done. I wrote more about that pattern in Release Checklist Azure DevOps: Standardize Readiness.
Conclusion
The fix for agent 5.279.0 is one line, but it deserves a decision rather than a reflex:
- The agent no longer mounts
/var/run/docker.sockinto Linux container jobs by default - Set
mapDockerSocket: trueonly on containers that truly need Docker - Prefer host steps or service containers where you can
- Use the escape hatch on self-hosted agents only with an expiry date
- Inventory your container jobs before the rest of your pool updates
If you want the next pipeline change to be a checklist item instead of an incident, add a pipeline review step to your release definition. Install the free Checklist Extension and make "agent and task changes reviewed" part of what Done means for every release.
Frequently asked questions
Reacties
Nog geen reacties. Wees de eerste om te reageren.
Plaats een reactie

Geschreven door Funs Janssen
Software Consultant
Ik ben Funs Janssen. Ik ontwikkel software en schrijf over de beslissingen die daarbij komen kijken: architectuur, ontwikkelmethoden, AI-tools en de zakelijke impact van technische keuzes. Deze blog is een verzameling praktische aantekeningen van echte projecten: wat schaalbaar is, wat misgaat en wat in blogvriendelijke voorbeelden vaak over het hoofd wordt gezien.
Inhoud


Use a release checklist Azure DevOps teams trust to standardize readiness, prevent incidents, and ship confidently every time.

