← christeceno.com

Design note · September 2026

How the software factory works

A curated task goes in. A reviewed draft pull request comes out. Nobody writes a prompt in between, a person is asked exactly twice, and the expensive model only touches the steps that need judgment. This is how it is put together and why.

The problem it solves

Small teams do not run out of ideas; they run out of hands. Work arrives as email, Slack threads, meeting transcripts, and rows on a task board, and most of it is noise. Turning the actionable fraction into shipped code has three bottlenecks: deciding what is worth doing, writing the code, and checking that the code is right. Agents are good at the middle one, bad at the first, and cannot be trusted alone with the third.

The factory is built around that asymmetry. It automates the middle, keeps a person on the first and third, and puts deterministic code, not a model, in charge of whether anything moves forward.

Intake: filter locally, decide by hand

Everything that could become work flows into one place: unread email, Slack mentions, video call transcripts, and the task boards. A model served on my own machine reads all of it, classifies each item, links it to the task it belongs to, and ranks what is left. This is high-volume work with a fixed output schema, which is exactly the shape a small local model handles well, and it costs nothing per call.

The output is a ranked inbox. Each candidate can be approved, dismissed with a reason, or snoozed. Approved candidates become runs. Dismissal reasons are kept, so the filter learns what I do not want to see again. That decision is deliberately manual: what a company should build next is not a question I want a model answering unsupervised.

The run: eight stages, four kinds of worker

Every approved task walks the same path. Each stage is done by whichever worker is cheapest and still correct: a local model, a cloud model, a person, or plain deterministic code.

  1. Intakelocal A classifier assigns one of four classes: build (code change with a testable outcome), investigate (a question about live data, answered in the task body, no code), spec (too thin to build from, so a proposed spec comes back for review), or human (I am the instrument: send X to Y, take the meeting). Only build continues down this path.
  2. Researchcloud The agent reads the repository and the live data and writes a plan into the task body: what will change, what will be measured, what it is unsure about.
  3. Gate 1human I approve the plan, or answer the questions it raised, or send it back. No code exists yet.
  4. Buildcloud The agent implements against the approved plan and nothing else.
  5. Verifytests Lint, unit tests, schema validation, and screenshot diffs where the change touches a page. No model is involved. A red check stops the run.
  6. Publishverified A draft pull request opens with a measured before and after in the body, and the task record is updated.
  7. Gate 2human I review the pull request. This is the only human touch between the plan and production.
  8. Releaseverified Merge, deploy, task moved to Done, memory updated with anything the run learned.

Why exactly two gates

A gate belongs where a wrong answer is expensive and a human is the cheapest way to catch it. Before any code exists, the plan is the whole cost of a mistake, and reading a plan takes two minutes. Before anything ships, the pull request is the whole cost, and reviewing a diff with green checks is fast. Everywhere else, a wrong answer is caught by a test or costs a few cents of retry, so a gate there would only add latency.

Gates are durable state, not prompts

An approval cannot live inside an agent session, because I might not look at it for hours. Each gate is a state the run parks in. The run resumes from that state when the decision lands, on any machine, with nothing lost. That single constraint shaped most of the architecture.

The gates are also where the dashboard earns its keep. It is the one decision surface: the inbox, the board, the strip of things waiting on me. Chat is notify-only. Decisions made in a chat thread are hard to find later; decisions made on a card with a reference are not.

Deterministic checks between every phase

The agents never decide whether their own work is good enough. Between every phase a fixed set of checks runs as code: linting, the test suite, schema validation of anything the agent produced as structured output, and pixel diffs for visual changes. A failure stops the run and surfaces the reason. This is what makes the two-gate design safe: by the time a pull request reaches me, the mechanical questions have been answered mechanically.

The agents themselves are tested the same way. Each has a binary evaluation suite, pass or fail, no sliding scale. A proposed change to an agent's instructions is tried as a single-file mutation, run against the suite, and kept only if nothing that passed before now fails. Prompt engineering becomes a build step with a test gate, like any other change.

Local versus cloud: measured, not assumed

The routing rule is simple: every step gets the cheapest thing that can do it correctly. The interesting question is what the local tier can actually carry, so I benchmarked candidate models against the two real prompts the local tier serves in production, a 25-item classification batch and task linking at the production ceiling of 150 candidates, on a 128 GB Apple Silicon machine running an MLX server.

ModelOn diskClassify batchTokens/sLink F1Holds the output schema
Gemma 4 26B MoE, 4-bit15 GB11.1 s850.970 of 2 at 150 items
Gemma 4 26B MoE, 8-bit26 GB18.2 s511.006 of 6
Gemma 4 31B dense, 4-bit17 GB51.9 s181.004 of 4
Gemma 4 31B dense, QAT 4-bit27 GB78.4 s121.003 of 3

MoE: mixture of experts, 26B parameters total with about 4B active per token. Measured September 2026 on the production prompts; public benchmark scores were not used.

Two things fell out of the numbers. First, the dense 31B builds score higher on public leaderboards and lose here anyway: a dense model reads roughly 15 GB of weights per token against about 4 GB for the 4B-active mixture, so it is bandwidth-bound at 12 to 18 tokens per second, and the extra intelligence is invisible on a task the smaller model already saturates. Second, the fastest build is not the safe one: the 4-bit mixture is the quickest but drifts off the required JSON shape at the 150-item ceiling, while the 8-bit holds it every time. So the 8-bit is the quality tier and the 4-bit the speed tier, with a smaller Ollama model as failover if the server is down.

Everything that needs judgment goes to Claude: research, implementation, review. Everything with a fixed schema and high volume stays local at zero marginal cost. Everything that can be a test or a script is a test or a script.

Memory

Runs accumulate knowledge: a trap in a codebase, a partner API that lies with a 200 status, a decision and why it was made. Each is written as a small file with its source cited, and mirrored into a vector store with locally computed embeddings so any later run can find it by meaning. Hazards are tagged so they inject themselves when their trigger words appear in a new task, without anyone remembering to look them up.

What it is not

It is not autonomous. It ships nothing I have not read twice, once as a plan and once as a diff. It does not pick its own work; it proposes. And it is not general: it knows one company's repositories, boards, and partners, which is precisely why it is useful there. The point was never to remove the human from the loop. It was to make sure the human's two decisions are the only two that need making.

Chris Teceno · christeceno.com · Resume