Contributing to pixtuoid
Thanks for your interest! PRs are welcome — especially new themes, sprite and
decoration polish, and Source adapters for agent CLIs we don’t support yet
(the agent CLIs plus the OpenClaw gateway already wired up are listed in the README).
Before you start, read AGENTS.md at the repo root (and the
nested crates/*/AGENTS.md for the crate you touch). It holds the load-bearing
architecture invariants and conventions. Many things that look like bugs are
documented, intentional design: read the whole item, its doc comment and the
comments on the lines it governs, before changing it.
Build & test
Requires a recent stable Rust toolchain and just
(brew install just). On Linux you also need lld, pkg-config and the ALSA
headers (apt install lld pkg-config libasound2-dev), and, with no GPU driver,
Mesa’s software Vulkan for the GPU tests (apt install mesa-vulkan-drivers).
The git hooks and most CI jobs call justfile recipes.
just # list recipes
just preflight # pre-push gate: lint → clippy; `just preflight full` adds hack → test (CI's Rust recipes)
just fmt # auto-format
just test # the whole suite, under cargo-nextest (`just setup-tools`)
just test -p <crate> <filter> # fast loop while iterating on one crate
Don’t expect clippy to warm
test’s build — its check-mode (rmeta) builds carry over only build scripts and proc-macros, so iterate with one of them.
Activate the git hooks once per clone: git config core.hooksPath .githooks
(pre-commit = just fmt-check; pre-push = just preflight, lint + clippy;
the tests are CI’s).
CI gates
CI is the gate. Beyond the tests and the feature powerset (just preflight full
runs those locally), it runs the jobs below; all but hygiene and zizmor’s
offline audits are invisible to preflight, so a green preflight does not mean a
green PR.
A PR’s pushes run only the light tier (every job without
if: inputs.full: linters, formatters, unit tests on every platform and
compile checks), and its ci-gate judges that tier. Both tiers run on the
merge queue’s draft PR (mergify/merge-queue/…), whose ci-gate the queue
merges on, batching up to .mergify.yml’s batch_size PRs in one run, and on
a push to main or a manual dispatch
(two-step CI). CodeQL and
CodSpeed skip drafts. The jobs:
- api-surface — committed
cargo public-apigoldens atapi/<crate>.txt; regenerate withjust api-surface+ commit when the public surface moves. - docs (
just doc-check) — rustdoc with-D warningsover private items, the bins, the examples and eachDOC_TARGETStriple, plus the doctests nextest skips. - generated drift (
just gen-readme-check gen-art-check gen-icons-check compare-selftest) — generated sprites, icons and README freshness, and the image comparator. - smoke · npm package generator (
just npm-check) — the release binaries and the hook shim’s silent exit, and the npm package generator + OpenClaw plugin contract. The README’s media drift (just gen-media-check) is reported in smoke as evidence, not a gate. - windows-check / windows-test — msvc cross-lint on every PR, and the full suite on a real Windows runner.
- other-unix-check (
just check-other-unix) — FreeBSD cross-lint for the other-unix arms. - wasm-check — builds the site’s wasm (
just gen-wasm) and caps its gzipped size (just gen-wasm-check). - site —
site.yml: format, lint, types, knip and unit tests on every push; the demo-reading test, e2e and Lighthouse on a build with freshly built wasm, so a Rust change that breaks a wasm export the page calls fails before it deploys. - snapshots —
cargo insta; fails on a pending OR orphan.snap, the rotjust testcan’t see. - hygiene — the same
just lintrecipes preflight runs (its CI job exists so a skipped local preflight can’t land a lint break), includingjust ci-observability(policy/ci-observability/: contracts for the silent, costly workflow failures actionlint and zizmor can’t see, and behavior tests of the workflows’ own shell) andjust fixture-pii(gitleaks over the committed capture tree); the capture-tree rules ridejust testinstead. - zizmor — workflow/action security: symbolic-or-SHA pins, credential-dropping checkouts, exact inline suppressions.
- One automatic Claude reviewer per
REVIEW.mdlens ridesclaude-readonly-review.yml: a read-only model job on the trusted default branch, the PR diff, title, body and the lens’s prior threads as inert data, and a separate least-privilege publisher that opens a review thread per finding and sets the lens’sclaude-review/<lens>status.claude.ymlrefuses fork PR heads. - CodeQL stays the advanced workflow (
codeql.yml): explicit languages, a SARIF health gate on Rust’snone-mode extraction, and an inline query filter droppingrust/cleartext-logging(WHY on the init step).
Releasing
Versioning
Pre-1.0: patch (0.y.Z) = bug fixes and polish only — no new public API,
nothing breaks. minor (0.Y.z) = everything else: new user-facing features
AND any breaking change to the published crates’ API. Both halves are machine-
applied on the release PR, not per-PR: release-plz derives the level from the
commit log (features_always_increment_minor in release-plz.toml is the
“features also bump minor” half), and release-plz’s own cargo-semver-checks
run is the “nothing breaks on a patch” half: a detected break raises the bump
to the next minor on its own, and the release PR’s body reports it. Never weaken
a lint to dodge the bump.
Cutting the release
release-plz owns every version number and the tag;
release.yml still owns every publish. Three steps, all human-initiated:
- Dispatch
release-plz.ymlfrom Actions, onmain. It openschore(release): vX.Y.Zfrom arelease-plz-*branch, with the workspace version, every path-dep requirement,Cargo.lockandCHANGELOG.mdrewritten. - Review it like any PR, but merge it by hand: the merge queue refuses it,
since its update would merge
mainin. Ifmainmoved, re-dispatch — never “Update branch”: only a dispatch recomputesCHANGELOG.mdfor the new commits, and the merge commit “Update branch” adds counts as a human’s, so the next dispatch closes this PR and opens a new number. Merged behindmain, or withmainmerged or rebased in, it failsrelease-plz.yml’srelease-mergeand publishes nothing. Raise the bump withcargo set-version --workspace X.Y.Z(cargo-edit) and push only for a breakcargo-semver-checkscannot see. A user-facing change that touched no packaged file (npm/,release.ymlpackaging) is not in the generated notes — add its line toCHANGELOG.mdby hand as the last commit before merging. - Merge it (squash). That merge is the irreversible step: the
releasejob publishes every crate to crates.io over OIDC, createsvX.Y.Zand a DRAFT GitHub release carrying the changelog; the tag then firesrelease.yml, which builds every target in its build matrix and the debs, attaches them, publishes the draft, and publishes the npm packages. The tag also starts a homebrew-core autobump.
The crates.io upload happens in the release job on the merge push, which
first waits for that commit’s own ci-gate: a failure or timeout publishes
nothing, and re-running the workflow after a passing ci.yml re-run resumes it.
The wait is in the workflow because the queue can land other PRs while a
release PR is open, so no merge-time check proves the tree it publishes.
cargo-semver-checks runs inside release-plz on the release PR, not as a CI
job; just semver reproduces its verdict locally.
Both jobs authenticate with RELEASE_PLZ_TOKEN, a fine-grained PAT scoped to
this repository with Contents and Pull requests read/write; release-plz.yml’s
header says why it cannot be the automatic token.
A release PR that release-plz closes and re-opens (it does that when the branch
carries non-bot commits) leaves a commit you pushed to it — a raised bump —
behind: git cherry-pick it onto the new branch. No committed frame carries the
version (BOARD_BRAND), so a release PR needs no just gen.
The tag also publishes outside this repo:
homebrew-core’s formula is autobump: true and builds from the tag tarball,
instantly, with DEFAULT features on macOS and Linux — the one configuration
our release never builds. Two consequences:
- A from-source build break lands in Homebrew’s CI, not ours. Anything
adding a system-library dependency needs a matching
depends_onin the core formula, in the same bump PR. - Their
test doblock is a public contract — see the “homebrew-core contract” comments atcrates/pixtuoid/src/app/sources_cli.rs,crates/pixtuoid-core/src/source/codex.rs. Change homebrew-core’stest dofirst, against the released version, so the next autobump stays green; the packaging-build action replays the block, so it changes in the same PR. Arelease-plz.ymldispatch opens no release PR while their block still calls a subcommand this tree dropped.
Do not try to preempt BrewTestBot: the formula is on homebrew-core’s
autobump list, so brew bump-formula-pr pixtuoid refuses by policy and the
bot opens the PR itself within ~3 hours of the tag. Watch THAT PR’s CI and
intervene only if it reds.
Publishing uses OIDC trusted publishing — CI carries no registry tokens; the per-crate/per-package Trusted Publishers must exist before the tag (#216).
The arc loop
Non-trivial work runs as an arc: design → build → gate → wrap.
- Pick — an issue (
gh issue list) or backlog item. - Grill the design — decide the open questions one at a time, each with a recommended answer, before writing code.
- Design gate (before build) — three lenses so slop dies in design: best-practice search (confirm the idiomatic way against real docs online, never memory) · adversarial design review (red-team the design before code exists) · deepening lens (would deleting this concentrate complexity or just move it? does the change deepen a module or add a shallow one?).
- Spec — synthesize into
docs/superpowers/specs/(LOCAL, git-ignored) and plan againstimpl-plan.prompt.md. - Mock gate (taste/visual work only) — ratify the AFTER visual before code
(
beautify-decorationskill). - Build — TDD: failing test → minimal impl → commit.
- Self-review — a standards+spec pass before pushing, INCLUDING the
whole-file comment audit: every file the PR touches — even by one line —
gets its entire comment population re-read against
AGENTS.md’s comment rules, and the cleanup rides the same PR (population and dispositions:REVIEW.md’s comment audit). Not the merge gate. - Merge gate — the gate; the
local-reviewskill runs its local rows; merging is@mergifyio queue, a release PR by hand. - Wrap — retro; a durable lesson becomes a mechanism (a test, a gate) or a line on the narrowest rule it amends — never an agent’s private memory, which nobody reviews and nothing executes.
Skills. Repo skills live in .claude/skills/
(committed; .agents/skills/ aliases them for Codex).
On a fresh machine or a non-Claude tool, git clone gives you the repo skills
and every just gate; this section IS the loop for tools without skills. Do
not scaffold a CONTEXT.md/docs/adr/ convention here — a declaration’s own
doc comment is the design record, and the nested AGENTS.md says only what its
crate IS.
The running order
| when | run |
|---|---|
| before code, if non-trivial (new seam / ≥3 files) | plan against impl-plan.prompt.md |
touched the --json / SourceStatus / OutcomeRow shape | just gen-contract |
| before push | nothing — the pre-push hook runs just preflight (never pipe it: a pipe eats the exit code) |
| while the work is in progress | push the branch with no PR: no workflow runs on a push to a branch other than main, so a PR-less branch costs the shared runners nothing |
| once the branch is ready to merge and a PR slot is free | open the PR ready: the light tier and the review bots run; a failure only the full tier catches surfaces in the queue, which dequeues the PR |
| when a REVIEW.md local row matches | the local-review skill |
| once the merge gate holds | @mergifyio queue |
| a source/lifecycle change | dogfood against live CC, or replay hermetically (tiers below) |
The e2e tiers live under scripts/lib/; none runs in CI. Cheapest first:
just openclaw-e2e (hermetic envelopes, free) · just replay <fixture> (a
captured rollout through the full headless path) · just openclaw-multi-e2e
(N real gateways, free) · just openclaw-backend-e2e (one BILLED turn) ·
just live-sources [id ...] (one BILLED turn per installed CLI; the only tier
proving a real CLI’s output becomes a sprite — sources with no invocation
entry are listed NOT COVERED, never skipped silently).
Advisory backstops that surface risk but never gate:
scripts/check_upstream_drift.py (wire-format drift) · just fixture-age
(which recorded fixtures a local CLI has moved past; LOCAL-only) ·
just bench / just bench-pacing / CodSpeed (local numbers authoritative; CI benches advisory).
Parallel sessions
- One
git worktreeand one cargo target per branch — a target shared across branches swaps uplifted examples and builds one branch’s types into another. Targets run to several GB each: checkdf -h /before parallel builds, and remove a PR’s worktree and local branch once it merges. - Fold before opening — a change to a surface an open PR already touches folds into it.
- The queue never idles — it checks one batch at a time (
.mergify.yml’smax_parallel_checks), so queue every PR that holds the gate, in priority order, at once; dequeue one only when it would jump a priority PR that is already green.
Conventions and architecture invariants
Both live in AGENTS.md (“Conventions”, “Architecture
invariants”), which every contributor and agent reads first.
Pull requests
- Review rules:
REVIEW.md. - AI-authored PRs get the
needs-human-verifylabel and a human visual check.
The merge gate
Green ci-gate; every lens bot’s required claude-review/<lens> status
success at the final head; every finding’s review thread resolved by its
disposition; zero open confirmed issue (blocking); each matching
local row’s run recorded as a PR comment starting
<!-- local-row:<row>:<head sha> -->, where <row> is the row’s first column
up to any colon or parenthesis, lowercased, each run of non-alphanumerics one
-, leading and trailing - dropped, and the sha is the head the run judged;
an update that only merges main in leaves the record standing. The
local-review skill runs those
rows. A published review passes whatever it found; a failed or missing status
is no review: comment /claude-review, else split the PR smaller.
Once the gate holds, comment @mergifyio queue (.mergify.yml):
entry is a command because no queue condition can confirm a finding or match a
local row. The queue tests up to batch_size PRs together on a draft PR
(mergify/merge-queue/…) running the full tier, then merges the PRs
themselves; it never updates a PR’s own branch
(batches: “the original PRs
are the ones merged”), so the bots’ statuses on the PR’s head are the ones its
queue_conditions read. media-regen.yml’s
bot PR (bot/media-regen, docs/images/ only) queues itself and needs no
generated-art record: it renders main’s merged code, which each look PR’s lens
already read as evidence.
The bots never review a fork PR on their own: a maintainer approves its CI
run, then comments /claude-review, again after every push. Its author can
resolve their own threads, so before merging read each thread’s resolvedBy
and its reply. Its bot verdict is advisory, since the
author can steer it through the diff, so the maintainer reads the diff too.
The bots skip Dependabot as an actor, so a maintainer comments it on its PRs
too, again after every Dependabot rebase.
Dispositions
Every finding reaches exactly one terminal state in its review thread: FIXED ·
REFUTED (cite the mechanism, per AGENTS.md; add one where none exists. Before
adding code for a finding, establish its case is reachable: when a test or
sweep shows it isn’t, that test is the mechanism and no defensive code lands) ·
RE-SCOPED → #N (real and INTRODUCED — or first made reachable — by this
change, and bigger than the PR: split it off into #N; a redesign that brings
the finding into scope ends FIXED) · FOLLOW-UP → #N (real and PRE-EXISTING,
or introduced, non-blocking, no bigger than the PR and found after round 2
per the convergence contract, and not FIXED in place
— in place fits a small defect inside code this change already touches,
adding no local row — so it is fixed in #N; a defect in another session’s
tree cites that session’s PR). A
disposition is the reply that resolves the thread, STARTING with its state:
FIXED: … · REFUTED: … — <mechanism> · RE-SCOPED → #N: … ·
FOLLOW-UP → #N: …, where #N is a PR other than this one: open, merged, or
closed under the open-PR cap with the fix on its branch. A
re-flag of an already-dispositioned finding replies with the original’s
disposition (link it). “Acknowledged” and “surfaced” are not states. Sweep at
the FINAL merge head; check WHICH commit a bot re-flag was raised against
before re-litigating.
Convergence contract
- Churn budget — a diff whose added + modified lines exceed ~1500 is split (stacked PRs) before review. Pure deletions are exempt once censused; a change that both adds and deletes at scale is two PRs.
- Deletion census — before deleting N members of a class, the full list and its criterion land in the first commit or the PR body (#943).
- Fixes are never put off; severity ends the loop. Rounds 1 and 2 each fold every accepted finding, whatever code it names, into ONE commit: a defect this change introduced is FIXED in it, since a clean-up left for later rarely lands (Google). After round 2 a fold takes only what the merge gate confirms blocks the merge, and a non-blocking finding is a FOLLOW-UP: a review moves on once only non-blocking suggestions remain (GitLab). A blocking issue confirmed in a fold STOPS the loop: revert the fold and re-land smaller, or re-scope.
- A fold’s behavior change ships a test that fails without it; the rest of a fold is a revert, a deletion, a comment or doc change, or a refactor the existing tests cover. Anything else reverts the fold.
- A fix round adds no new gate — a wanted check is its own PR, asserting
facts in its own layer (a Rust fact from Rust, never a Python regex over
.rs).
Handy gh commands
gh pr checks --watch # live CI status
gh issue develop <number> --checkout # branch linked to an issue
gh run rerun --failed # rerun only failed CI jobs
Adding a new agent CLI
The registration steps (4–7, 9) and step 12’s roster literals are test-forced —
skipping one fails just test. Step 8 is forced only for hook-only sources;
step 10 by the theme guards; steps 1–3, 11 and step 12’s #[test] are on you.
- Verify the wire format against the CLI’s actual source/releases first —
transcript location, line shape, hooks, session identity; pin every fact
to an upstream file/version. Audit its HOME RESOLVER per axis in the
same pass — PROBE the installed artifact rather than trusting docs; an
unmirrored axis is fail-silent: the watcher polls a directory the CLI
never writes and the office stays empty (#880). Resolver axes are
deliberately NOT drift-watched — re-run the probe matrix when the CLI majors.
A custom root gets ONE
pub fn <cli>_home(), called by both the watcher’sdefault_paths()and the installer’sdefault_config_path()so they can’t disagree. - Write the source module —
crates/pixtuoid-core/src/source/<name>.rs:SOURCE_NAME, aLineDecoderfn (one JSONL line →Vec<AgentEvent>), a label deriver, unit tests per event mapping. Format knowledge lives HERE. - Implement the
Sourcetrait (an asyncrun(self, tx)watching + decoding until the session universe ends). Hook-only CLI? Skip the decoder, trait, and step 7:transcript: Nonein the registry row, format knowledge in ahook.customdecoder (it must claim EVERY event), and do step 8 instead. - Add ONE
SourceDescriptorrow insource/registry.rs— label prefix, decoder, hook keying,tool_id_key(verify against a CAPTURED tool call, not a neighbour — kimi’sToolCallcost a source its tool ids), truthful capability flags,verified_version+version_probe. Lifecycle policy derives from the flags; you do not edit the reducer. - The descriptor’s
nameis the roster —registered_source_names()projectsREGISTRY, and the conformance suite then requires a fixture. Thesources --jsongolden (crates/pixtuoid/tests/snapshots/cli/sources.json) must list it:SNAPSHOTS=overwrite just test -p pixtuoid --test cli_json. - Record the fixture — the test steps in
crates/pixtuoid-core/tests/AGENTS.md(a RECORDED SessionStart scenario viajust capture-fixture), thencargo insta review. - Wire it into
runtime/driver.rs::build_source_set(the one construction site; the registry drives the guard test, not the spawning). - If the CLI has hooks, add an
install/target (aTargetrow +merge_install/merge_uninstall+ averify_schemafn mirroring the target’s own config format + the registered-events↔decoder-arms guard). - Add a row to
site/src/sources.json(status,featured, per-OSplatforms), thenjust gen-readme. Pinned toregistered_source_names()bysupported_sources_manifest.rs. - Add the per-source badge hue — a
SourceColorsfield + value in EVERY theme file +badge_colorin the manifest row; the coverage, legibility and site-bridge tests fail until it exists. - Drift-watch in the same PR: a
check_upstream_drift.pyrow where one is owed — which surfaces owe one issource/drift.rs’s header, read it there. A row is four steps: the const, theinsertin that crate’ssrc/drift_surface.rs,just gen-drift-surface(commit both fragments), and theSURFACE_ROWSrow plus its selftest case (the case census fails without it). - Three roster literals in three test binaries (a scoped run misses
them): the row-by-row byte pin in
corpus_check.rs;TOOL_ID_KEY_UNPROVENintests/sources/captures.rs; a case row +#[test]incrates/pixtuoid/tests/wire_to_pixels.rs.
License
By contributing, you agree your contributions are licensed under the same terms as the project.