Star 历史趋势
数据来源: GitHub API · 生成自 Stargazers.cn
README.md

CodeGraff workshop rats in coral coats

macOS · Linux · Windows One binary, 3.7 MB Zero dependencies Built in Zig 0.17 dev

justrach/codegraff | Trendshift

curl -fsSL https://github.com/justrach/codegraff/releases/latest/download/install.sh | sh

Download CodeGraff for Mac
Apple Silicon · macOS 14+ · signed and notarized · everything included

Evaluated on FrontierHarness tasks. Graff includes a reproducible FrontierHarness evaluation runner, with recorded outcomes and explicit protocol differences. See how we measure it for the public task suite, grading, and the limits of comparisons with the published board.

The desktop app

The desktop app uses Electron and embedded Chromium, with native macOS window controls and SwiftUI Activity and computer-use panels. Coding continues in graff acp; the window is a client of the same harness used by the terminal.

CodeGraff's bright desktop, framed in rice paper with workshop artwork

Chat tabs, searchable model selection, effort and fast controls, collapsed tool activity, and explicit working/finished/interrupted states keep the conversation readable. Appearance includes White, Black, Website and the official CodeGraff palette. Mention $gui-theme or @gui-theme in the GUI to create a custom theme.

CodeGraff Agents panel with local peers and a handoff request, framed in coral with the workshop crew

Agents brings Graff-to-Graff coordination into the GUI. See sessions in the current workspace or across the laptop, their published tasks and activity, and recent messages. Select a peer in the panel's composer to send a message or handoff request. Delivery is queued to the recipient's next step; browsing history does not consume their inbox. The optional profiler records anonymous per-agent resource measurements, with identities and message contents excluded from feedback exports. See the Agents guide.

CodeGraff's dark Changes panel beside the conversation, framed in cobalt with a rat reviewing a proof

Changes shows local staged, unstaged and untracked edits, diffs, worktrees and recent commits. Drag its divider to give the review more room. The browser pane renders directly in Chromium and supports navigation, find, zoom and element pins. Optional macOS computer use exposes native app inspection and input after the user enables it and grants the operating system permissions.

Presentation frames pair CodeGraff workshop artwork with unchanged GUI captures and scripted demonstration content. Click a desktop image for the full-size UI. No private conversation or workspace data is included.

Build and launch from source on Apple Silicon macOS 14+ with Bun, Zig and Xcode command-line tools installed:

./script/build_and_run.sh

The development bundle contains the production UI, Chromium, Bun, graff and the native bridge, and starts its own local server. It does not need a separate dev server or Kuri. Source builds use local development signing. The downloadable release is Developer ID signed and notarized. Installed releases from v0.0.291 download updates in the background and wait for Restart to update. Earlier versions need one manual replacement in Applications to enable the updater.

Profile and test without a model. The Performance menu and desktop profiler tool record bounded, local measurement reports. Startup paint timing, streaming responsiveness, process resources and acceleration status are measured separately. No reports are uploaded automatically. From apps/native:

bun run build
bun run test:desktop
bun run test:visual
bun run test:performance

The visual and performance scenarios use production GUI components with scripted inputs and block engine/model API calls. See the desktop guide and visual test guide for scope and limitations.

What can I ask it?

If you could do it at a computer, you can ask graff to do it for you:

  • "Build me a little app to track my workouts." It writes it, runs it, and shows you.
  • "Turn this folder of messy CSVs into one clean spreadsheet."
  • "Figure out why my site is slow, then fix it."
  • "Scrape these five pages and summarize them."
  • "Run an experiment: try three versions of this and tell me which scores best."

It works in your real terminal, on your real files, with the real internet, and it can spin up a team of sub-agents in parallel.

Don't write code? You don't have to. Say what you want in plain English.

Same model, fewer tokens

Same grok-4.6, same SuperGrok seat, same tasks. graff vs grok-build vs OpenCode. Lower is better on every named axis. Full tables: graff-evals/hillclimb/baseline.md.

12 shipped-PR fixtures (run-20260901-121759-composite on 284, --suite inhouse):

harnesspasswallcallstokenslist$RSS
graff12/12220s53234k$0.328.7M
grok-build12/12490s601.12M$1.07155M
OpenCode12/12235s77675k$0.681.0G

Graff is the unique frontier on pass, wall, calls, tokens, list$, and RSS on this remasure. (First-token is not scored — graff's 0.0s is a boot mark, not first model SSE. RSS is ReleaseSafe process peak.)

On the 3-task spine (exact-reply + file-ops + fix-fib) graff was 19.9s / 8 calls / $0.048 vs grok 32.3s / 8 / $0.147 and OpenCode 31.2s / 8 / $0.101.

Graff carries context through three steps: reuse the stable setup, run small programs over the working context, and return focused results. That keeps the next step supplied with useful information while reducing repeated input.

A workshop rat examines a proof: stable context, small programs, and focused results keep useful context in the harness

Install

Desktop (Apple Silicon, macOS 14+). Download CodeGraff v0.0.291, quit any running Codegraff copies, open the disk image, and drag Codegraff.app onto Applications. Eject the disk image and open Codegraff from Applications. The notarized bundle includes Graff, Chromium, Bun and the native macOS components; you do not need a separate CLI, runtime, developer tools or local server. Verify the download checksum. For a graff command in your terminal, use the CLI installer below.

Desktop builds from v0.0.291 check for updates online and download them in the background. Choose Restart to update when your work is finished, or use Codegraff → Check for Updates…. Automatic downloads can be disabled in that menu. Earlier builds need one manual installation to enable the updater. An app update replaces the bundled Graff engine together with the interface. A CLI installed separately by the command below has its own update lifecycle; that command downloads a CLI archive, not the notarized desktop installer.

CLI (macOS · Linux · Windows).

curl -fsSL https://github.com/justrach/codegraff/releases/latest/download/install.sh | sh

From a checkout: ./install.sh (binary in ~/bin; HARNESS_NO_PATH=1 skips PATH edits). Windows: unpack graff-*-windows.tar.gz from the latest release and put graff.exe on PATH.

graff login                     # free codegraff key
graff login kimi                # Kimi Code OAuth
graff login codex               # ChatGPT / Codex (or reuse ~/.codex/auth.json)
graff key set deepseek sk-...   # any other provider
graff                           # REPL
graff --model grok-4.6          # pin a model
graff -p "how many TODOs in src/?"

graff acp is the Agent Client Protocol spawn (Zed External Agents). Recipe: docs/acp-registry.md.

Why it's small

metricmeasured
binary~3 MB, zero runtime deps
cold start~1.8 ms
full agentic turn~12 MB peak RSS
8 parallel subagents+0.4 MB each
fat tool outputone 4 KB handle, whatever the result's size

Same model, same endpoint, the older Rust codegraff used 4.3× the memory and ~14× the disk for a dead-heat turn. Method: architecture.md.

Script it

from harness_sdk import Harness
with Harness(yolo=True, model="gpt-5.5") as h:
    print(h.ask("what is 2+2?"))
import { runAgent } from "@codegraff/sdk";
for await (const ev of runAgent({ prompt: "summarize README.md", yolo: true })) {
  if (ev.type === "text") process.stdout.write(ev.text);
}

graff --json / graff --schema generate the SDKs (sdk/). Remote: graff serve. Embedders: --no-local-tools + a sandbox MCP — Embedding graff.

CLI, slash commands, providers, permissions
graff [flags]                 REPL
graff -p "prompt"             one-shot (answer on stdout)
graff login [codegraff|codex|kimi]
graff key set <provider> <key>
graff mcp add <name> -- <cmd>
graff learn <command>
graff --schema

--model <name>   --yolo   --json   --no-local-tools
--subagent-model <name>   --max-model-calls N

One-shot has no human at the gate: pre-approve in .harness/settings.json or pass --yolo. Full flag list: graff --help. Learning: docs/local-learning.md. Skills: docs/skills.md.

/model /models /clear /new /goal /loop /review /never
/plan /yolo /strict /effort /compact /rewind /btw
/skills /plugins /mcp /save /resume /sessions /help

Bare / is a filterable menu. Esc interrupts the turn. /help is the live catalog.

modewhat it does
defaultask before writes, MCP, and non-read-only bash
--yolo / /yoloskip every prompt (CI, -p)
/planread-only explore
/strictevery message is a tool

Providers: Anthropic, OpenAI, DeepSeek, xAI, Z.AI, Kimi, Codex (ChatGPT login), Vercel, OpenRouter, MiniMax, Xiaomi, Groq, Cerebras, Mistral, plus one workspace router in .graff/.config.router. graff models refresh pulls catalogs. Claude-subscription OAuth is deliberately not supported.

How we measure it

Two layers, under graff-evals/. They answer different questions and neither of them is a leaderboard claim.

Layer 1 — the in-house runner (run.py, harnesses.json, tasks/). Every task is one JSON file: fixture files, a prompt, and a deterministic shell check that decides pass/fail inside a materialized sandbox. Held-out checks live in hidden/ and are injected through $TASK_ROOT after the harness exits, so the agent never sees them. Most harnesses take --model, so the same task set can be driven through different harnesses on one model, and each run records wall time, first-output latency, peak RSS, CPU and token usage alongside the verdict, as JSONL plus a summary table.

45 tasks in five suites — core (12, sequential single-file work), rlm (5, scatter-gather across files), swe (6, multi-file bugfixes), mcp (10, a fixture MCP bench), inhouse (12, bug shapes distilled from shipped PRs). --suite all is core+rlm+swe; mcp and inhouse are opt-in. 25 harness configurations are declared, covering this project's variants plus several other CLI agents. A task that requires a capability a harness lacks is skipped, not scored as a failure. Cost is recomputed from tokens at published list rates, because a flat-rate subscription prints $0.0000 and that is a plan, not a price.

What this layer proves: that a change moved a measured number on a fixed, deterministic task set. What it does not prove: anything about the live repo — the inhouse fixtures are distilled shapes, not the codebase.

Layer 2 — frontier-harness/. It runs the same 30 tasks as FrontierHarness Eval — 21 from Terminal-Bench 2.1 and 9 from DeepSWE — in Docker, under a protocol that is deliberately not the same bench seat (see "What these runs are not" below, and PROTOCOL.md). The board side is a pinned snapshot of the published results, not a live query. TB tasks are graded by running the public tests/test_outputs.py inside the task container after the agent exits — pass is pytest exit 0. The 9 DeepSWE tasks the upstream pack treats as having a hidden grader are scored out of band by grade_swe.py against the tests datacurve-ai/deep-swe actually ships, using the same images and the same prepare/test.sh protocol, reading the verifier's reward.json. A missing reward.json is recorded as FAIL, never inferred. A competing agent is run locally on the same images and the same tests.

What these runs are not

  • Not same seat as the published board. The later recorded runs used an eval-only system-prompt append (BENCH_APPEND, passed as --append-system-prompt). It is task-shaped coaching the board's harnesses did not get. It never touched the shipped prompt in prompt_text.zig, and an appended-prompt result must not be placed next to a board result as a peer. The honest number is the un-appended first pass.
  • Different model. The published board is Kimi K3; the recorded runs are mostly a different model. To compare fairly: empty BENCH_APPEND, same model, TB-21 only, and say so.
  • Different runtime. The official eval restores a prepared VM. We docker run the public image and, on stripped images, add a CA bundle and install pytest so TLS and the tests can run at all. That is infrastructure, not a hint, but it is not bit-identical.
  • Asymmetric cost columns. The locally run competing agent logged no token events, so its list price is missing — a telemetry gap, not zero. It is also driven through its own CLI and its own runner, so it shares the images and the tests but not the harness path. The chart refuses to place a row with no cost data on the frontier.
  • Mixed-model harness rows are a different comparison. Entries that run another agent on its own native default model are not points in a same-model series, and mcp is always run in one mode because the other is a different tool catalog.
  • Some recorded misses are environmental — an agent wall-clock cap, a server that did not outlive the agent process, a leftover build artifact breaking a file-layout constraint — and are written up as such in FAILURES.md. On the DeepSWE side apply_failed is not excused: it is a real failure.

Reproduce

cd graff-evals

./run.py --harness graff                       # core+rlm+swe; mcp/inhouse are opt-in
./run.py --harness graff,grok --model grok-4.6 # harness-vs-harness, same model
./run.py --harness grok --task fix-fib --reps 3
zig build && ./run.py --harness graff-dev      # the locally built binary
./run.py --interactive                         # pick a task, watch it live

Results land in results/run-<stamp>.jsonl; .sandboxes/ keeps the last run's working directories for post-mortems. Both are disposable.

cd graff-evals/frontier-harness
export FH_GRAFF_MODEL=grok-4.6   # or kimi-k3 + MOONSHOT_API_KEY

python3 fh_run.py --suite tb  -j 2 --fresh --out results.jsonl      # TB-21
python3 fh_run.py --suite swe -j 2 --out swe-results.jsonl          # DeepSWE patches
python3 grade_swe.py grok-4.6                                       # grade those patches
python3 plot_tb21.py

The competing agent has its own runner, fh_exo.py, and its own binary (EXO_BIN); fh_run.py does not drive it.

This layer is not turnkey. It needs Docker, a Linux build of the binary, the upstream task pack, the pinned terminal-bench tests and a clone of datacurve-ai/deep-swe, staged where the scripts expect them — PROTOCOL.md has the locations. Model selection is an environment variable. No credential is committed here: the metered path reads its key from the environment, and the subscription path copies an existing local credentials file into the task container.

Working on codegraff

scripts/install-hooks.sh          # once
scripts/eval-tier1.sh             # offline, ~20s warm
python3 scripts/eval-tier2.py     # model-backed, opt-in

Tier 1 is zig fmt, the 600-line ceiling, test reachability, zig build test (suite count never shrinks), named goal/loop/todo invariants, and SDK drift. Docs-only pushes skip it. In-house PR fixtures: graff-evals/ (--suite inhouse).

License

Modified GNU AGPL-3.0 (LICENSE). Network use triggers Section 13. Authors Rach Pradhan (justrach) and Yu Xi Lim (yxlyx) reserve the right to offer proprietary or hosted versions. A recipient's AGPL licence is perpetual unless they breach it. Commercial permission without copyleft exists only if both authors grant it jointly in writing, and is revocable.

Built in Zig 0.17 dev · AGPL-3.0 (modified) · architecture.md · CHANGELOG · uxlog.md

关于 About

graff — a fast agentic coding harness in Zig: multi-provider, MCP, workflows, DGM evolution loop, TS/Python SDKs

语言 Languages

Zig63.1%
TypeScript16.0%
Python11.9%
Rust2.7%
JavaScript2.6%
Swift1.0%
Lean0.8%
CSS0.7%
Shell0.6%
HTML0.5%
C0.0%

提交活跃度 Commit Activity

代码提交热力图
过去 52 周的开发活跃度
1938
Total Commits
峰值: 302次/周
Less
More

核心贡献者 Contributors