Skip to main content

Code Mode pipelines

Declare the steps a Code Mode session runs in .schemabounce/pipeline.yaml: planning, checks, reviews and approvals, and which of them can block a push.

A pipeline is the ordered list of steps a Code Mode session runs in your repository. You write it in .schemabounce/pipeline.yaml. The platform runs the steps in order and decides whether the work may be pushed. The agent cannot declare a gate passed.

Four kinds of step exist. Everything else is configuration of those four.

KindWhat it doesCan it block the push?
turnThe agent works in a role you describe and writes named artifacts.No
checkRuns shell commands in the work container. Pass or fail.Yes, when a command is blocking
reviewGrades the work against criteria with thresholds.Only a fresh review
approvalParks the session until a person approves, edits or rejects.Yes, always

A repository with no .schemabounce/ folder gets one step, implement, and nothing gates the push. Add a pipeline when you want planning, tests or sign-off to be part of every session.

A minimal pipeline​

apiVersion: schemabounce.com/v1
kind: RepoPipeline
pipeline:
- uses: implement
- name: tests
uses: check
commands:
- { name: unit, run: npm test }

The agent does the work, then tests runs npm test. If it fails, the agent gets the output in a fix turn and tests runs again, up to 2 times. After that the session ends failed.

Every step has a unique name: lowercase letters, digits, dashes and underscores, 40 characters at most. A step without a name is named after what it uses (implement, then implement-2).

A fuller pipeline​

This one plans twice, designs the screens, asks a person to sign off, then implements each planned unit with its own check. Three whole-repository checks and two reviews follow.

apiVersion: schemabounce.com/v1
kind: RepoPipeline
setup:
- { name: install, run: npm ci }
pipeline:
- name: architecture
uses: plan
guidance: Decide the approach in a few sentences. Name the modules you will touch.
- name: ui-design
uses: platform/designer
with: { output: screens }
reads: [architecture]
- name: breakdown
uses: plan
reads: [architecture, ui-design]
maxUnits: 6
- name: sign-off
uses: approval
reads: [architecture, ui-design, breakdown]
onReject: { goto: breakdown }
- forEach: breakdown.units
steps:
- uses: implement
- name: unit-check
uses: check
commands:
- { name: unit-tests, run: npm test -- --related }
- name: types
uses: check
commands:
- { name: tsc, run: npx tsc --noEmit }
onFail:
retry: { maxAttempts: 3 }
then: hold
- name: lint
uses: check
commands:
- { name: eslint, run: npx eslint . }
- { name: prettier, run: npx prettier --check ., severity: advisory }
- name: e2e
uses: check
commands:
- { name: smoke, run: npx playwright test smoke, timeout: 15m }
artifacts: [playwright-report]
- name: self-review
uses: review
context: shared
criteria:
- { name: matches-ticket, threshold: 0.7 }
- name: final-review
uses: review
context: fresh
criteria:
- { name: correctness, threshold: 0.8 }
- { name: tests-cover-change, threshold: 0.7, blocking: false }
complete:
endpoint: pr

What the keys do:

  • setup runs once before the first step. A failed setup stops the session.
  • reads names earlier steps whose artifacts this step receives. An approval must read something.
  • forEach: breakdown.units repeats its steps once per unit of an earlier plan step. It cannot nest.
  • when: { paths: [...] } skips a step unless the session changed matching files.
  • optional: true keeps a step out of the gate and out of the honesty rule below.
  • budget: { time: 15m, spendUsd: 2 } caps one step inside the session's own cap.
  • artifacts lists files a check keeps as evidence.
  • complete.endpoint is branch (push the branch, the default) or pr (open a pull request). merge is refused.

In a pipeline, a command blocks by default. Mark one severity: advisory to report it without failing the step.

Built-in steps and uses:​

uses: accepts three kinds of reference.

ReferenceResolves to
turn, check, review, approvalThe primitive itself. Configure it inline.
plan, implement, platform/designerA platform built-in.
repo/<name>.schemabounce/steps/<name>.yaml in your repository.

plan is a read-only turn that splits the ticket into units. implement does the work, or the current unit inside a forEach. platform/designer is a read-only turn that describes the screens a change touches; its output parameter says what to design. Anything else is refused with the list of valid names.

Only these sources exist today. Sharing steps across a workspace and from the marketplace is planned and not available yet.

A step definition file​

A definition is a reusable step. The file name must match name, and kind is StepDefinition.

# .schemabounce/steps/api-designer.yaml
apiVersion: schemabounce.com/v1
kind: StepDefinition
name: api-designer
primitive: turn
description: Describe the API surface before any code is written.
readOnly: true
with:
surface:
default: HTTP endpoints
description: What to design.
role: |
You design the {{ with.surface }} this ticket needs. List each one with its
request, response and error cases. Do not edit files.
produces:
- { name: design, path: design.md, schema: design/v1 }
# .schemabounce/steps/security-review.yaml
apiVersion: schemabounce.com/v1
kind: StepDefinition
name: security-review
primitive: review
context: fresh
criteria:
- { name: no-secrets-logged, threshold: 0.9 }
- { name: input-validated, threshold: 0.8 }

Use them from the pipeline:

apiVersion: schemabounce.com/v1
kind: RepoPipeline
pipeline:
- name: api-design
uses: repo/api-designer
with: { surface: REST endpoints }
- name: build
uses: implement
- name: tests
uses: check
commands:
- { name: unit, run: go test ./... }
- name: security
uses: repo/security-review

Passing a with: key the definition does not declare is an error, as is leaving out a parameter marked required: true. A step can add commands or criteria to its definition, but not reuse a name the definition already has.

A definition is text and settings. It narrows what the agent may do (tools.deny, tools.ask) and never widens it. It cannot carry a plugin.

When a step fails​

A failing check or review does not end the session at once. The default:

  1. The agent gets the failing output or critique in a fix turn.
  2. The same step runs again.
  3. After 2 attempts (the default), the step's then applies.

Change this with onFail:

- name: e2e
uses: check
commands:
- { name: smoke, run: npx playwright test smoke }
onFail:
retry: { goto: build, maxAttempts: 3 }
then: hold
  • retry.maxAttempts sets the attempts. The platform ceiling is 5; a higher number is clamped.
  • retry.goto replays the pipeline from an earlier turn in the same block instead of fixing in place. It must name a turn step.
  • then is what happens when the attempts run out. fail ends the session as failed (you can continue it). hold parks the session for the owner to decide. continue reports the failure and moves on.

A step that gates the push cannot choose continue. fail is the default for a gating step. A shared review defaults to continue.

For an approval, onReject: { goto: <step> } sends the person's note back to an earlier step.

Reviews: shared and fresh​

context decides who judges.

shared (default)fresh
Where it runsInside the session, in the agent's own contextIn an isolated reviewer, separate from the session
Blocks the pushNeverYes
Can prove a plan criterionNoYes
Use it forAdvice the agent can act onThe verdict you trust before a push

A verdict produced inside the coder's sandbox is not trustworthy, so a shared review only advises.

Choosing the reviewer​

A fresh review can name the engine and model that grade the work, for example Codex reviewing code Claude Code wrote.

- name: final-review
uses: review
context: fresh
engine: codex
model: your-model-id
criteria:
- { name: correctness, threshold: 0.8 }
  • engine is claude-code or codex. gemini-cli is not available yet and is refused.
  • model is required with engine. The platform never picks a model for another engine's reviewer.
  • A model without an engine is refused. A shared review cannot name an engine.
  • With no engine, the review uses the session's own engine, model and credential.

A named reviewer is paid for in this order: the session owner's own connected license for that provider, then workspace credits (billed as the session's spend), then the work waits for the owner. A SchemaBounce-held key is never used.

What a fresh review does today​

The isolated reviewer runs only where the platform holds a reviewer isolation certification. Hosted SchemaBounce does not hold one yet, so today a fresh review does not grade the diff:

  • If the session changed no files, the step passes.
  • Otherwise the session parks and the owner gets a push approval for the exact diff. Approve it and the pipeline goes on. Deny it and the session ends failed with review_rejected.

Once the certification is in place, the reviewer runs on the engine and model you declared. The platform reads them from the pipeline it stored when the session started, so nothing inside the session can change which reviewer grades it.

Plans with proof​

A plan step writes a plan. In a pipeline, each acceptance criterion in the plan names the evidence that proves it. The platform computes "done" from that evidence instead of taking the agent's word.

{
"schema": "plan/v2",
"summary": "Retry failed webhook deliveries with backoff",
"units": [
{
"id": "retry",
"title": "Retry with backoff",
"intent": "Retry a failed delivery 3 times, doubling the wait each time, so a short outage does not drop events.",
"files": ["internal/webhooks/retry.go"],
"acceptance": [
{ "text": "A failed delivery is retried 3 times", "proof": "check:unit-check/unit-tests" },
{
"text": "The wait doubles between retries",
"proof": "check:unit-check/backoff-timing",
"command": { "name": "backoff-timing", "run": "npm test -- backoff", "proposed": true }
},
{ "text": "No secret appears in retry logs", "proof": "review:final-review" }
]
},
{
"id": "docs",
"title": "Document retries",
"intent": "Describe the retry schedule in the webhooks page.",
"docsOnly": true,
"acceptance": [{ "text": "The page states the schedule", "proof": "review:final-review" }]
}
]
}

A proof is one of:

  • check:<step> proves the criterion when the whole check step passes.
  • check:<step>/<command> narrows it to one command of that check.
  • review:<step> cites a fresh review.

The planner is given the list of proofs your pipeline can produce. It can also propose a new command for a check ("proposed": true). A proposed command runs in the work container like any repository command, and the approval card marks it as proposed.

Why a plan is refused​

The planner gets the reason back and tries again. A plan is refused when:

  • A proof names a step or command the pipeline does not have.
  • A proof cites a shared review.
  • A unit has no criterion proven by a check, so nothing runs the code. Mark the unit "docsOnly": true if it changes no behavior, and it may rest on reviews alone.
  • A unit has no acceptance criteria, or more than 8.
  • The plan has more units than the plan step's maxUnits (default 12, ceiling 40).

The proof matrix​

The session page shows every criterion with its status:

StatusMeaning
ProvenThe cited step passed. The evidence links to its recorded result.
FailedThe cited step or command failed.
Waiting on youThe step is held, skipped, or a person's override has no approval on file.
Not run yetThe step has not run, or the code changed after it ran.

A per-unit check keeps its verdict when a later unit changes the code. That criterion shows Proven on an earlier commit. A top-level check counts only on the code as it stands: if the agent changes code after it passed, it runs again before the push.

Review the plan first​

The composer has a Review the plan first switch. It makes the session stop after planning so you approve the plan before any code is written.

  • If the pipeline already has an approval step that reads a plan and comes before any step that writes code, the switch uses it. A step marked optional is turned on.
  • Otherwise the platform adds an approval step named review-plan after the last plan step, reading the plan steps before it.
  • A pipeline with no plan step has nothing to review, and the switch does nothing.

On the approval card you can edit the plan: change units, criteria and proofs. Approve with my edits continues the work from your version. Send back returns the plan to the agent with a note. The card shows problems (an unknown proof, for example) before you can approve.

At least one blocking check​

A declared pipeline.yaml must contain at least one non-optional check with a blocking command, anywhere in the pipeline including inside a forEach. Without one, nothing runs the code before a push, and the file is refused:

pipeline: the pipeline has no non-optional check with a blocking command,
so nothing runs the code before a push (dec-041). Add a check step

A check whose only commands are severity: advisory does not count.

Repositories with workflow.yaml and validate.yaml​

A repository with only the older files keeps working. The platform converts them into a pipeline when a session starts:

  • Plan phases in workflow.yaml become plan steps before implement.
  • The gate in validate.yaml becomes one validate check after implement, retrying with the gate's maxFixAttempts.
  • Review phases become fresh reviews. They already held the work for the owner, so nothing gets less strict.

If pipeline.yaml exists, it wins and the older files are ignored for the steps.

Note that validate.yaml requires severity: blocking or advisory on every step. pipeline.yaml defaults to blocking.

See the effective pipeline​

Open a session's page and choose the Definition tab. It lists the steps in the order they run, after defaults, limits and legacy conversion, and shows each step's result as the session goes. A session keeps the pipeline it started with, so editing pipeline.yaml later does not change a running session.

To check a repository before you start a session, use the readiness grader described in Code Mode on your machine.

Limits​

Going over a limit is an error for structure and a clamp with a warning for counts. These come from the platform.

LimitValue
Steps in one pipeline, nested steps included32
Step definition files in .schemabounce/steps/32
Commands in one check16
Setup commands8
Criteria in one review12
reads on one step16
with parameters on one step or definition16
Artifact globs on one check8
Fix attempts5 (default 2)
Plan units40 (default 12)
Acceptance criteria per plan unit8
One command's timeout5 seconds to 30 minutes (default 10 minutes)
Definition role text8,000 characters
Step guidance text2,048 characters
One .schemabounce/ file256 KB

Common questions​

Can the agent skip a check?​

No. The platform stores the pipeline when the session starts and lets a push through only when every gating step in it has passed. The runner reports results. It cannot mark a gate passed.

Why does my shared review never block anything?​

A shared review runs in the agent's own context, so its verdict is not independent. It advises and never gates. Set context: fresh for a review that blocks the push.

Can I use workspace/... or market/... steps?​

Not yet. Today a step uses a primitive, a built-in, or a repo/<name> definition from your repository.

Continue exploring