Showing posts with label mirrord. Show all posts
Showing posts with label mirrord. Show all posts

Sunday, September 27, 2026

The shell script passed. The cluster test could not deploy the intercept.

The shell script passed because a person was holding the CLI. The automated test had no CLI to hold. That is the bug tekton-dag found while moving pull-request intercepts off the laptop and into the cluster.

It turned up inside a larger experiment. I am running eight or more projects at once, and Obsidian is the shared record between the Cursor sessions doing that work. Most of the time the record is there so a session can take the next bounded item, finish it, and leave the queue understandable. This one was the other use. A problem popped up on one project, and the same notes had to carry a deep fix without losing the rest.

A pull request builds the changed app and deploys it as a PR pod next to the live stack. Traffic that carries x-dev-session: pr-N must be served by that pod. Traffic that does not carry the header must stay on the baseline, untouched. The pipeline is supposed to install that split itself and then prove it. For a long time the install step was a Telepresence intercept, and the proof was a script that exited zero.

Two ways of running the same test

From a laptop, the path is a handful of commands. Install the traffic manager, run telepresence intercept, name the deployment, set the header, and point the intercepted traffic at a local process. The person running the script has an admin kubeconfig. The CLI is on their PATH. Those runs intercepted. That is what the session notes remembered as "E2E with Telepresence intercepts passing."

The refactor put the same job inside Tekton. A task pod in the cluster was supposed to deploy the intercept with nobody at a keyboard. The sidecar we gave it was ghcr.io/telepresenceio/tel2:2.20.0. That image is the traffic agent. It contains one binary, traffic. telepresence is not in it. telepresence helm install failed with not found, the container exited, and kubectl wait still succeeded because the pod was briefly Ready during a ten-second sleep.

You cannot deploy this intercept without the CLI commands. Connect, create the intercept, set the header filter, and aim it at the PR workload are all telepresence subcommands. The in-cluster image cannot run them. The automated test was not a stricter version of the shell script. It was a different program, missing the only tool that does the work.

The check on top of that was warn-only. It printed NOT CONFIRMED and exited 0. Automated runs on main stayed green through the middle of September. On 26 September the check was changed to fail when a hop was not confirmed. The next run on main went red. Telepresence had not changed. The test had stopped agreeing with a log line.

Why copying the CLI into the pod was not the fix

We did try to put a real Telepresence CLI in the cluster. A client pod with the CLI, its own service account, a tun device, and a privileged namespace still could not hold the intercept the way a laptop does.

  • A newer CLI could not create intercepts against the traffic manager already in the cluster. The client and the manager have to match closely.
  • Once a matching CLI did create an intercept, header-matched requests hung. The client image had no iptables, so the session that carries stolen traffic back to the handler never started.
  • A whole-port intercept, which does not need the header filter, stole baseline traffic too. That breaks the contract.
  • Injecting the traffic agent restarts the live deployment. The laptop workflow does not do that to the running stack in the same way, because the intercepted process is on the laptop.

The net finding is to take Telepresence out of the in-cluster tests. It remains a fine laptop tool: install it on a workstation, run the CLI, debug against the cluster. It does not belong in a pipeline that has to deploy the intercept from a pod. Pull request 109 removes the in-cluster task, the traffic-manager install from the runner, and the Telepresence leg of the test matrix.

What the two paths actually were

Shell script, from a laptop. This one intercepted.
Laptop
telepresence CLI
admin kubeconfig
header filter set by a command
→ Traffic manager
in the cluster
takes the CLI's intercept
→ Split
header → local process
no header → baseline
Automated job, inside the cluster. This one could not deploy.
Tekton task pod
image tel2:2.20.0
binary: traffic
no telepresence CLI
→ No command to run
cannot connect
cannot create the intercept
cannot set the header
→ Nothing served
PR pod was a sleep loop
check printed NOT CONFIRMED
exit code 0

Why the manual run did not become the CI run

The manual test and the CI job shared a name. They did not share a procedure. A person at a laptop was doing five things the pipeline never did, and each one was load-bearing. Drop any of them and the intercept is a different test.

What the person supplied What the CI job had instead
The telepresence CLI on a laptop PATH, and the commands that create the intercept. A pod image whose only binary is traffic. The commands the script ran were not in the container.
An admin kubeconfig. The CLI could install, connect, and edit the live workload. A namespace service account. The same binary, run as that account, is forbidden from the calls the laptop made.
A real process on the laptop, which answered the stolen request. A PR pod whose command was while true; do echo "PR build"; sleep 30; done. It could not answer HTTP.
Eyes on the full terminal. A failed connect was the next line of output. An artifact collector that ran kubectl logs without asking for the whole log. The default is ten lines per container, so the error a person would have read was cut off.
A pass the person could see: the request landed on their process, and a request without the header did not. A pass the script could exit: print NOT CONFIRMED, then exit 0. Green meant the task ended, not that the pod served the request.

That is why months of green CI did not contradict a passing shell script, and did not confirm one either. The script was evidence about a laptop session. CI was evidence that a different program returned zero. Wiring the script into GitHub Actions would not have closed the gap, because the script's success still depended on the person: their binary, their credentials, and their reading of the log.

The same split showed up again after the pipeline moved to mirrord, which can run its CLI in a pod. A hand test in the cluster routed correctly. The pipeline run of that same idea did not. Two details that a person fills in without noticing:

  • The account. The hand test used an admin user. The proxy pod used the namespace default service account, which is not allowed to create the agent or to open the port-forward the agent connection needs. The deploy step only waited for Ready. A container with no readiness probe is Ready the instant it starts, including one that exits a second later on Forbidden. The hand test never hit that account, so it never hit that failure.
  • The header string. mirrord matches the header as name: value, colon then a space. The hand test used that form. The task passed x-dev-session:pr-N with no space. The filter could not match, every request stayed on the baseline, and a check that only looked for the header echoed back in the response still printed pass. Both pods echo the header. Echo is not identity.
A manual pass proves the tool in the hands of the person who ran it. It does not prove the job. CI has to run the same command, as the same kind of account, against a process that can answer, and fail when the PR pod's own log does not show the matched request and only the matched request.

How MetalBear fixed the in-cluster path

MetalBear ships mirrord as a CLI whose job is to run a process in the context of the cluster. The CLI does not have to live on a laptop. The tool image for the pipeline contains it. The task starts a proxy pod, and that pod runs mirrord exec.

mirrord creates an agent pod beside the live deployment. The agent is what steals traffic. Steal mode with an HTTP header filter takes only the requests whose header matches x-dev-session: pr-N and hands them to the process mirrord is wrapping. In the pipeline that process is a small relay to the PR pod, which is running the built image as the real app. Requests that do not match stay on the baseline pods. The live deployment is not restarted to install a sidecar, and the client pod does not need a tun device or iptables.

In-cluster intercept with mirrord. The CLI runs in the pod. The agent does the split.
1. Tekton task starts a proxy pod as its own service account, not as the namespace default. The pod's image includes the mirrord CLI.
2. mirrord exec in that pod creates the agent next to the target deployment and connects to it through the Kubernetes API. The filter is (?i)^x-dev-session: pr-N$. mirrord matches the header as name: value, colon then space. A filter written without the space never matches.
3. Agent, steal + header filter. Matching request → relay → PR pod (the built app). Any other request → baseline pods, unchanged.
4. Proof, from the PR pod's access log. A marked URL with the header must show up there. The same mark without the header must not. Remove the intercept and the first check goes red. Steal every request and the second check goes red.
                incoming request
                       │
                       ▼
              mirrord agent (in cluster)
              header filter on the Service
                    │            │
         x-dev-session          no match
           : pr-N                  │
              │                    ▼
              ▼              baseline pods
         proxy pod            (live stack)
         mirrord CLI
              │
              ▼  relay
           PR pod
     (built image, real app)
              │
              ▼
     access log must contain
     the matched probe only

The first honest pass of that proof is what the shell scripts never had to show. A response that echoes pr-N is not evidence: the baseline echoes the same header. Pod Ready is not evidence either: a container with no readiness probe is Ready the instant it starts, including a process that crashed a second later. The access log is the pod admitting it saw the request. The unmatched probe is the control that says it did not see the other one.

mirrord is in the test matrix because the pipeline can deploy it. The CLI is in the image, the agent is a pod the task is allowed to create, and the filter lives in config the task writes. Telepresence's in-cluster image could not issue the commands the shell script depended on.

Lessons

None of these showed up as a failing assertion the first time. Each one showed up as a story that was already believed.

  • A manual pass is about the person who ran it. Their binary, their credentials, the string they typed, and the log they read. The job has to be given those same ingredients or it is a different test.
  • A green pipeline is only as honest as its weakest check. Exit 0 on NOT CONFIRMED is not a result. Pod Ready is not a result. A response that echoes the header is not a result, because the baseline echoes it too. The check needs a negative control that fails when the feature is absent.
  • Keep the log you will debug from. kubectl logs on a selector keeps ten lines per container unless you ask for the rest. The line that explained the failure was in the part that was thrown away.
  • Do not match the word ERROR and call it a failure. A healthy tool prints ERROR for noise. The gate has to name the fatal message. A gate that matches the noise stops the run that was about to show you the real bug.
  • Find the bug on a cluster you can still see. Four failures in a row were each debugged from a log zip after CI had deleted the cluster. Every one of them was a two-minute question with kubectl on a live local cluster. CI is the clean confirmation after that, not the place the diagnosis happens.
  • Write down the path you closed, and why. "E2E passed" was already in the notes from the spring. Without a newer note that says that sentence was wrong, the next session will treat it as the fact and spend another afternoon on it.

The notes were how the next session knew where to start

The experiment has two speeds, and I wanted to know whether one record could serve both.

The fast speed is a batch. Across the other projects the work is a list of bounded items: one check, one hole, one small change, then stop. A session should be able to open a project, read what is left, do that item, and write down where it stopped. Nobody re-briefs it. The vault is the queue. That is the automation. Not a scheduler inside the notes, and not the notes launching the next session. Cursor does the work. Obsidian is what lets the next session start from the item instead of from a blank chat, so eight projects can move without eight briefings.

The slow speed is a problem that will not fit in an item. This intercept was that. It did not fit in one chat. Through the day the story changed: both backends were green, then I rejected a whole-port intercept, then both greens were false, then the in-cluster Telepresence path was closed, then mirrord's own green was a race. A later Cursor session does not remember that. It remembers the last thing someone typed, or an older note that says the test passed. If the notes only knew how to hand out the next batch item, this fix would have been a fresh investigation every time I came back to it, and the other projects would have stalled while I held the context in my head.

The record that survived is in Obsidian, written and read through obsidian-mcp. Three notes, and they do different jobs:

  • Project state says what done means right now. Merge only when a request with the session header shows up in the PR pod's log, and a request without it does not. A green that cannot show that is not done.
  • A working note is the timeline of what we believed against what the artifacts showed. February's "E2E passing" sits in the same table as the sleep loop and the exit code. When the next session starts, it reads the correction before it reads the claim.
  • A decision freezes the choice and the alternatives that were rejected. mirrord is the in-cluster backend. Telepresence stays on the laptop. Do not bring the in-cluster path back without the same routing proof. Debug on the local cluster; let CI confirm.

Cursor calls that vault before the work, not after a wrong turn. get_project_context returns the state, the latest session, and the decisions. search_memory finds the decision by the question you actually have, instead of by which chat happens to be open. When a choice should outlive the session, record_decision writes it as its own note. When the session stops, capture_work_session appends what changed and what the next person should not redo.

That is what kept the solution moving. The next session did not "fix" the red build by letting NOT CONFIRMED exit 0 again, because the state file said the earlier green was empty. It did not spend another round of CI on an in-cluster Telepresence sidecar, because the decision already recorded why that path was closed. It took the next red run to the local cluster, because the decision said a log zip is the wrong tool for a question kubectl can answer while the pod still exists.

The notes were also wrong on the same night, and that is part of the lesson. An early write said mirrord had met the contract. An audit a little later showed its green was the same kind of lie: the proxy died on a forbidden call, Ready flickered, and the response echoed a header the baseline would have echoed anyway. The working note was corrected in place. A vault that cannot be corrected just launders the last confident sentence, which is how "E2E passing" survived from the spring until the check was made fail-closed.

The same notes do both jobs. They hand the next bounded item to a session that is working through a batch, and they hold a deep fix when a problem pops up, including the paths already closed. The notes did not find this bug. The artifacts did. What the experiment showed is that I did not have to drop the other projects to stay inside this one.

Obsidian is the scratch pad. The repo ledger is the lesson.

Something about the relationship to SDLC-SPDD is coming into focus, and I do not want to pretend it was designed this way on day one.

SDLC-SPDD still guides. It says which phase the work is in, what done means, and what is allowed to become a permanent record. It does not have to hold the intermediate work. That gets pushed to Obsidian: the scratch pad a session writes while the work is still messy. Project state, the timeline of what we believed, the decision, the correction an hour later, the CI run that is still red, the guess that turned out to be wrong. That accumulation is the point. A deep fix needs the noise, because the noise is where the last wrong sentence is written down and then crossed out. A batch needs it too, because "where this item stopped" is not a lesson yet. It is a handoff.

That pile should not become the memory of the repository. A later change in tekton-dag does not need the night's chronology, the ten-line log truncation, or the hour where both backends looked green. It needs the sentence that survived the night. A manual pass does not prove the job. A green check needs a negative control. Do not throw away the log you will debug from. Do not fail a healthy tool for printing ERROR. Find the bug on a cluster you can still see, and let CI confirm. Write down the path you closed.

Those sentences are what come back. The guide stays in charge of the return trip. A capture during the work is staged and stays out of git. At the end of the work, an accept step promotes what is worth keeping into the committed lessons ledger. Nothing else is supposed to land there, and the ledger is not edited by hand to make a story fit. The scratch pad can be wrong at 9pm and corrected at 11. The ledger should only receive the correction that still matters after the noise has been thrown away. Intermediate work lives in Obsidian. SDLC-SPDD keeps guiding, and it keeps the lesson.

SDLC-SPDD guides. Intermediate work goes to Obsidian. Only a distilled lesson comes back.
1. SDLC-SPDD guides
Which phase this is. What done means. What is allowed to become a permanent record. The intermediate notes do not go here.
↓ intermediate work pushed out, noise included
2. Obsidian — the scratch pad
Project state, the believed-versus-true timeline, the decision, the handoff ("stopped here"). This pile stays here. It is not the memory of the repository.
Stays in the vault
chronology
false starts
CI run still red
a sentence corrected at 11pm
→ 3. Distill
one lesson that is still true after the noise is thrown away
a manual pass does not prove the job
↓ staged, not committed yet
4. Accept
Review the staged lesson. Promote it, or leave it. The scratch pad is not copied across.
↓ committed
5. SDLC-SPDD lessons ledger
The sentence the next change can retrieve. Not the night. Not edited by hand to make a story fit.
The guide stays. The scratch pad holds the middle. The ledger receives the lesson. Intermediate work is pushed to Obsidian so the repository does not fill up with the night. What comes back is small enough that the next change can retrieve it without replaying that night.

What I will trust a green run to mean

A green intercept job has to print that the PR pod served the matched request and did not serve the unmatched one. Local kind produced that line. The intercept job on pull request 109 then produced it on a clean cluster, and that pull request is merged. A laptop script from last spring is not what made it done.

If you are wiring the same kind of test: run the command your script runs from inside the job that is supposed to replace the script. If that command is not in the image, you are not automating the test. You are automating the exit code.

The change is tekton-dag pull request 109. mirrord's own model is documented at mirrord.dev.