Code Mode pipelines
Declare the steps a Code Mode session runs in .schemabounce/pipeline.yaml: planning, checks, reviews and approvals, and which of them can block a push.
A pipeline is the ordered list of steps a Code Mode session runs in your repository. You write it in .schemabounce/pipeline.yaml. The platform runs the steps in order and decides whether the work may be pushed. The agent cannot declare a gate passed.
Four kinds of step exist. Everything else is configuration of those four.
| Kind | What it does | Can it block the push? |
|---|---|---|
turn | The agent works in a role you describe and writes named artifacts. | No |
check | Runs shell commands in the work container. Pass or fail. | Yes, when a command is blocking |
review | Grades the work against criteria with thresholds. | Only a fresh review |
approval | Parks the session until a person approves, edits or rejects. | Yes, always |
A repository with no .schemabounce/ folder gets one step, implement, and nothing gates the push. Add a pipeline when you want planning, tests or sign-off to be part of every session.
A minimal pipeline
apiVersion: schemabounce.com/v1
kind: RepoPipeline
pipeline:
- uses: implement
- name: tests
uses: check
commands:
- { name: unit, run: npm test }
The agent does the work, then tests runs npm test. If it fails, the agent gets the output in a fix turn and tests runs again, up to 2 times. After that the session ends failed.
Every step has a unique name: lowercase letters, digits, dashes and underscores, 40 characters at most. A step without a name is named after what it uses (implement, then implement-2).
A fuller pipeline
This one plans twice, designs the screens, asks a person to sign off, then implements each planned unit with its own check. Three whole-repository checks and two reviews follow.
apiVersion: schemabounce.com/v1
kind: RepoPipeline
setup:
- { name: install, run: npm ci }
pipeline:
- name: architecture
uses: plan
guidance: Decide the approach in a few sentences. Name the modules you will touch.
- name: ui-design
uses: platform/designer
with: { output: screens }
reads: [architecture]
- name: breakdown
uses: plan
reads: [architecture, ui-design]
maxUnits: 6
- name: sign-off
uses: approval
reads: [architecture, ui-design, breakdown]
onReject: { goto: breakdown }
- forEach: breakdown.units
steps:
- uses: implement
- name: unit-check
uses: check
commands:
- { name: unit-tests, run: npm test -- --related }
- name: types
uses: check
commands:
- { name: tsc, run: npx tsc --noEmit }
onFail:
retry: { maxAttempts: 3 }
then: hold
- name: lint
uses: check
commands:
- { name: eslint, run: npx eslint . }
- { name: prettier, run: npx prettier --check ., severity: advisory }
- name: e2e
uses: check
commands:
- { name: smoke, run: npx playwright test smoke, timeout: 15m }
artifacts: [playwright-report]
- name: self-review
uses: review
context: shared
criteria:
- { name: matches-ticket, threshold: 0.7 }
- name: final-review
uses: review
context: fresh
criteria:
- { name: correctness, threshold: 0.8 }
- { name: tests-cover-change, threshold: 0.7, blocking: false }
complete:
endpoint: pr
What the keys do:
setupruns once before the first step. A failed setup stops the session.readsnames earlier steps whose artifacts this step receives. An approval must read something.forEach: breakdown.unitsrepeats its steps once per unit of an earlier plan step. It cannot nest.when: { paths: [...] }skips a step unless the session changed matching files.optional: truekeeps a step out of the gate and out of the honesty rule below.budget: { time: 15m, spendUsd: 2 }caps one step inside the session's own cap.artifactslists files a check keeps as evidence.complete.endpointisbranch(push the branch, the default) orpr(open a pull request).mergeis refused.
In a pipeline, a command blocks by default. Mark one severity: advisory to report it without failing the step.
Built-in steps and uses:
uses: accepts three kinds of reference.
| Reference | Resolves to |
|---|---|
turn, check, review, approval | The primitive itself. Configure it inline. |
plan, implement, platform/designer | A platform built-in. |
repo/<name> | .schemabounce/steps/<name>.yaml in your repository. |
plan is a read-only turn that splits the ticket into units. implement does the work, or the current unit inside a forEach. platform/designer is a read-only turn that describes the screens a change touches; its output parameter says what to design. Anything else is refused with the list of valid names.
Only these sources exist today. Sharing steps across a workspace and from the marketplace is planned and not available yet.
A step definition file
A definition is a reusable step. The file name must match name, and kind is StepDefinition.
# .schemabounce/steps/api-designer.yaml
apiVersion: schemabounce.com/v1
kind: StepDefinition
name: api-designer
primitive: turn
description: Describe the API surface before any code is written.
readOnly: true
with:
surface:
default: HTTP endpoints
description: What to design.
role: |
You design the {{ with.surface }} this ticket needs. List each one with its
request, response and error cases. Do not edit files.
produces:
- { name: design, path: design.md, schema: design/v1 }
# .schemabounce/steps/security-review.yaml
apiVersion: schemabounce.com/v1
kind: StepDefinition
name: security-review
primitive: review
context: fresh
criteria:
- { name: no-secrets-logged, threshold: 0.9 }
- { name: input-validated, threshold: 0.8 }
Use them from the pipeline:
apiVersion: schemabounce.com/v1
kind: RepoPipeline
pipeline:
- name: api-design
uses: repo/api-designer
with: { surface: REST endpoints }
- name: build
uses: implement
- name: tests
uses: check
commands:
- { name: unit, run: go test ./... }
- name: security
uses: repo/security-review
Passing a with: key the definition does not declare is an error, as is leaving out a parameter marked required: true. A step can add commands or criteria to its definition, but not reuse a name the definition already has.
A definition is text and settings. It narrows what the agent may do (tools.deny, tools.ask) and never widens it. It cannot carry a plugin.
When a step fails
A failing check or review does not end the session at once. The default:
- The agent gets the failing output or critique in a fix turn.
- The same step runs again.
- After 2 attempts (the default), the step's
thenapplies.
Change this with onFail:
- name: e2e
uses: check
commands:
- { name: smoke, run: npx playwright test smoke }
onFail:
retry: { goto: build, maxAttempts: 3 }
then: hold
retry.maxAttemptssets the attempts. The platform ceiling is 5; a higher number is clamped.retry.gotoreplays the pipeline from an earlier turn in the same block instead of fixing in place. It must name a turn step.thenis what happens when the attempts run out.failends the session as failed (you can continue it).holdparks the session for the owner to decide.continuereports the failure and moves on.
A step that gates the push cannot choose continue. fail is the default for a gating step. A shared review defaults to continue.
For an approval, onReject: { goto: <step> } sends the person's note back to an earlier step.
Reviews: shared and fresh
context decides who judges.
shared (default) | fresh | |
|---|---|---|
| Where it runs | Inside the session, in the agent's own context | In an isolated reviewer, separate from the session |
| Blocks the push | Never | Yes |
| Can prove a plan criterion | No | Yes |
| Use it for | Advice the agent can act on | The verdict you trust before a push |
A verdict produced inside the coder's sandbox is not trustworthy, so a shared review only advises.
Choosing the reviewer
A fresh review can name the engine and model that grade the work, for example Codex reviewing code Claude Code wrote.
- name: final-review
uses: review
context: fresh
engine: codex
model: your-model-id
criteria:
- { name: correctness, threshold: 0.8 }
engineisclaude-codeorcodex.gemini-cliis not available yet and is refused.modelis required withengine. The platform never picks a model for another engine's reviewer.- A model without an engine is refused. A shared review cannot name an engine.
- With no
engine, the review uses the session's own engine, model and credential.
A named reviewer is paid for in this order: the session owner's own connected license for that provider, then workspace credits (billed as the session's spend), then the work waits for the owner. A SchemaBounce-held key is never used.
What a fresh review does today
The isolated reviewer runs only where the platform holds a reviewer isolation certification. Hosted SchemaBounce does not hold one yet, so today a fresh review does not grade the diff:
- If the session changed no files, the step passes.
- Otherwise the session parks and the owner gets a push approval for the exact diff. Approve it and the pipeline goes on. Deny it and the session ends failed with
review_rejected.
Once the certification is in place, the reviewer runs on the engine and model you declared. The platform reads them from the pipeline it stored when the session started, so nothing inside the session can change which reviewer grades it.
Plans with proof
A plan step writes a plan. In a pipeline, each acceptance criterion in the plan names the evidence that proves it. The platform computes "done" from that evidence instead of taking the agent's word.
{
"schema": "plan/v2",
"summary": "Retry failed webhook deliveries with backoff",
"units": [
{
"id": "retry",
"title": "Retry with backoff",
"intent": "Retry a failed delivery 3 times, doubling the wait each time, so a short outage does not drop events.",
"files": ["internal/webhooks/retry.go"],
"acceptance": [
{ "text": "A failed delivery is retried 3 times", "proof": "check:unit-check/unit-tests" },
{
"text": "The wait doubles between retries",
"proof": "check:unit-check/backoff-timing",
"command": { "name": "backoff-timing", "run": "npm test -- backoff", "proposed": true }
},
{ "text": "No secret appears in retry logs", "proof": "review:final-review" }
]
},
{
"id": "docs",
"title": "Document retries",
"intent": "Describe the retry schedule in the webhooks page.",
"docsOnly": true,
"acceptance": [{ "text": "The page states the schedule", "proof": "review:final-review" }]
}
]
}
A proof is one of:
check:<step>proves the criterion when the whole check step passes.check:<step>/<command>narrows it to one command of that check.review:<step>cites a fresh review.
The planner is given the list of proofs your pipeline can produce. It can also propose a new command for a check ("proposed": true). A proposed command runs in the work container like any repository command, and the approval card marks it as proposed.
Why a plan is refused
The planner gets the reason back and tries again. A plan is refused when:
- A proof names a step or command the pipeline does not have.
- A proof cites a shared review.
- A unit has no criterion proven by a check, so nothing runs the code. Mark the unit
"docsOnly": trueif it changes no behavior, and it may rest on reviews alone. - A unit has no acceptance criteria, or more than 8.
- The plan has more units than the plan step's
maxUnits(default 12, ceiling 40).
The proof matrix
The session page shows every criterion with its status:
| Status | Meaning |
|---|---|
| Proven | The cited step passed. The evidence links to its recorded result. |
| Failed | The cited step or command failed. |
| Waiting on you | The step is held, skipped, or a person's override has no approval on file. |
| Not run yet | The step has not run, or the code changed after it ran. |
A per-unit check keeps its verdict when a later unit changes the code. That criterion shows Proven on an earlier commit. A top-level check counts only on the code as it stands: if the agent changes code after it passed, it runs again before the push.
Review the plan first
The composer has a Review the plan first switch. It makes the session stop after planning so you approve the plan before any code is written.
- If the pipeline already has an approval step that reads a plan and comes before any step that writes code, the switch uses it. A step marked
optionalis turned on. - Otherwise the platform adds an approval step named
review-planafter the last plan step, reading the plan steps before it. - A pipeline with no plan step has nothing to review, and the switch does nothing.
On the approval card you can edit the plan: change units, criteria and proofs. Approve with my edits continues the work from your version. Send back returns the plan to the agent with a note. The card shows problems (an unknown proof, for example) before you can approve.
At least one blocking check
A declared pipeline.yaml must contain at least one non-optional check with a blocking command, anywhere in the pipeline including inside a forEach. Without one, nothing runs the code before a push, and the file is refused:
pipeline: the pipeline has no non-optional check with a blocking command,
so nothing runs the code before a push (dec-041). Add a check step
A check whose only commands are severity: advisory does not count.
Repositories with workflow.yaml and validate.yaml
A repository with only the older files keeps working. The platform converts them into a pipeline when a session starts:
- Plan phases in
workflow.yamlbecomeplansteps beforeimplement. - The gate in
validate.yamlbecomes onevalidatecheck afterimplement, retrying with the gate'smaxFixAttempts. - Review phases become fresh reviews. They already held the work for the owner, so nothing gets less strict.
If pipeline.yaml exists, it wins and the older files are ignored for the steps.
Note that validate.yaml requires severity: blocking or advisory on every step. pipeline.yaml defaults to blocking.
See the effective pipeline
Open a session's page and choose the Definition tab. It lists the steps in the order they run, after defaults, limits and legacy conversion, and shows each step's result as the session goes. A session keeps the pipeline it started with, so editing pipeline.yaml later does not change a running session.
To check a repository before you start a session, use the readiness grader described in Code Mode on your machine.
Limits
Going over a limit is an error for structure and a clamp with a warning for counts. These come from the platform.
| Limit | Value |
|---|---|
| Steps in one pipeline, nested steps included | 32 |
Step definition files in .schemabounce/steps/ | 32 |
| Commands in one check | 16 |
| Setup commands | 8 |
| Criteria in one review | 12 |
reads on one step | 16 |
with parameters on one step or definition | 16 |
| Artifact globs on one check | 8 |
| Fix attempts | 5 (default 2) |
| Plan units | 40 (default 12) |
| Acceptance criteria per plan unit | 8 |
| One command's timeout | 5 seconds to 30 minutes (default 10 minutes) |
Definition role text | 8,000 characters |
Step guidance text | 2,048 characters |
One .schemabounce/ file | 256 KB |
Common questions
Can the agent skip a check?
No. The platform stores the pipeline when the session starts and lets a push through only when every gating step in it has passed. The runner reports results. It cannot mark a gate passed.
Why does my shared review never block anything?
A shared review runs in the agent's own context, so its verdict is not independent. It advises and never gates. Set context: fresh for a review that blocks the push.
Can I use workspace/... or market/... steps?
Not yet. Today a step uses a primitive, a built-in, or a repo/<name> definition from your repository.
Related Documentation
Continue exploring