Daniel SchwarzResearch Engineer, Polaris · application

A scroll-driven film · scroll to fly

An ordinary evening.

My agents are at work on the monorepo. X is open on the side.

Then X recommended this post.

A post from Tamay Besiroglu, quoting Google DeepMind’s Polaris announcement. Open on X ↗

Click.

Coding agents. Evaluation. "How you think and build." Everything I do every day suddenly had a name.

Let me show you what I've built.

A scroll-driven film · scroll to fly

I come from practice.
I work with agents all day.

I run 29 AI coding agents on a 19-app monorepo I call Terminal. This film shows what exists, the real numbers, what comes next, and why it led me to evals.

Terminal is the name of my own monorepo. It is not related to Terminal-Bench.

Three files hold the world together.

.gitmodules pins 25 submodules. index.config registers 19 apps with one layout. .gitignore keeps runtime state and secrets out of git.

19 apps. Walled in. One owner each.

Every app is its own repository with hard walls and exactly one owning agent. No agent edits another agent's app.

Every include has the same anatomy.

A template sets one structure for all apps. Backend is split by language and domain, and so is the database. Frontend is split into web, iOS, a shared protocol layer and API snapshots. Then come architecture, development infrastructure, production infrastructure and end-to-end flows. The largest app shows it best, here as structure only.

"Done" means a verified SHA.

Owners report the commit they shipped. The parent fetches it, verifies the SHA, and only then moves its pin. 1,206 of the parent's 2,342 commits are exactly that: pin moves.

One registry wires every stack to life.

Registry, port map and one resolver boot all apps at once on one machine. Adding an app means one registry entry, one port and one start script. Nothing else, and it never guesses.

One slot per machine. Same contract everywhere.

user/ holds a slot for each machine: Machine 1 (the master), Machine 2, Machine 3 and a cloud slot. Each slot has a schema-checked operator context, so product code never knows which machine it runs on.

Secrets travel sealed.

Apps hold typed references only. The Vault materializes them as files at start-up. Agents never get a master key.

.cursor is how I work with AI.

One meta repo with 955 commits steers every agent. It holds 44 rules (20 workspace-wide, 24 scoped to one include) and 23 skills. It also holds 779 plans (714 completed, 65 pending), 34 decision records indexed per include, hooks that block shell commands working off master, and 1,144 refinery findings. The workspace MCP config is kept empty on purpose; servers are connected per user.

libs: the most valuable include.

97 packages (56 TypeScript, 36 Python, 3 C#, 1 C++, 1 PowerShell), with 609 commits of their own. They grow only when the refinery finds real duplicates. Host code stays home. One path, always.

The rule behind it: build concrete first, abstract when a second case appears. No shared base type until two similar things exist.

Four top agents defend their domains.

Terminal owns pins and wiring. CEO sets direction but never invents roadmaps. Cursor Master enforces rules and hooks. Libs Master keeps libs minimal and necessary. Under them sit app orchestrators and a reviewer, 29 agents in total, and architecture moves only through my GO.

A factory, measured in lines, not commits.

Across the parent and 21 included repos: 2.77M lines added, 1.11M deleted, 1.65M net, and 18,949 files touched. 40% of everything written was later rewritten or removed. The parent repo alone has 2,342 commits on 87 active days.

Source and docs only, from git numstat. Excluded: vendored, generated and data directories, lockfiles, minified files, binaries and data formats (JSON, CSV, logs). Oct* = through 10 Oct.

How it got here.

$ git log --since=2026-05-27 | monthly
May   ▏                         4
Jun   ▍                        59
Jul   ▍                        67
Aug   █▏                      193
Sep   ████████████████████    1,410
Oct*  ████████▋                 609
$ ls .cursor/plans   714 completed · 65 pending

next, in order        status          todos open
Plan 1  Vault identity  pending · started   3
Plan 2  pipeline        pending             20
Plan 3  rule switch     pending             29

Commits: parent repo. Plans: .cursor/plans. Oct* = through 10 Oct. Plan rows: todo counts from the plan files.

Where I am right now

What comes next.

Everything above runs today. I build concrete first and abstract only when a second case appears. Early on, big changes straight on master were the fastest way to grow. Now the codebase is mature enough that precision matters more than speed. So the next step is three plans, run in order 1 → 2 → 3.

Every machine, no collisions.

The goal: agents work from every machine and from the cloud without colliding in the same tree. Secrets live in Vault. Only reviewed work lands on origin/master.

pending · in progress

Plan 1: a Vault identity plane.

A small, hardened VM next to production stores the secrets. Agents put and get. Every call names its location, person and process. Terminal issues the identity, and Vault runs its own auth check instead of OIDC.

pending

Plan 2: one worktree per agent.

Once identity exists, each agent gets its own worktree or cloud snapshot and a temporary branch that expires. Branch prefixes name the slot: machine-1/, machine-2/, laptop/, cloud/.

pending

Only reviewed work reaches master.

A reviewer agent checks every change, and only reviewed work reaches origin/master. A janitor removes branches whose work never finished. origin/master stays the only persistent state.

pending

Plan 3: then the rules flip.

Today's rules and locks say the opposite: work directly on master, on this machine. Plan 3 switches rules, hooks, AGENTS.md and replica scripts to the pipeline. It comes last on purpose: without Plan 2 it would only reword the old rule.

Where I am, exactly.

plan 1 Vault identity plane: pending, being started (3 todos open)
plan 2 worktree + branch pipeline: pending (20 todos open)
plan 3 rule switch: pending (29 todos open)

An umbrella plan (completed, 4/4) only links the three and fixes the order 1 → 2 → 3. Counts from .cursor/plans, 10 Oct.

The rebuild, in six rules.

ephemeral       short-lived credentials and environments per run
idempotent      setup can be rerun safely, any time
stateless       every run pulls what it needs fresh, from Vault
isolated        one worktree and one branch per agent
least privilege an agent gets only what its task needs
one truth       one registry, one secret source, one master

I chose these for agents to work safely. They are also exactly what a reproducible, gradeable agent run needs. That is an eval.

And then, agents fail.

Problems I observed in my own repos. Library drift: hosts written against a newer library API than the one they pinned. Pin lag: "done" reported, but the parent never moved its pin. Secrets at start-up: an app that wouldn't start because a secret wasn't wired. Tests that depend on where they run.

The cliffhanger: evals.

Every failure above passed the agent's own checks. The agent said "done", its tests were green, and the real state disagreed. A grader can't trust the transcript; it has to check the state the agent leaves behind.

After the remaining rebuilds, Terminal's next step is working with evals. I haven't studied them in depth yet. But I arrived here naturally, from practice, and I know exactly why they are needed. The pipeline I just showed is already a grader at heart: nothing counts as done until something other than the agent has checked it.

No coincidence.

This post was recommended to me while I was scrolling X between agent runs. I arrived at evals on my own, from practice, and I know exactly why they matter. What I need now is to learn how they are built and perfected, and that is exactly this job.

Seeing this post was no coincidence. It's my way into a new world. I'm ready to walk this path, and I hope to hear from you.

GitHub access on request · d.schwarz-bsnss@protonmail.com
Remote from Germany (CET) · built with my agent fleet, reviewed by me