Point Workbench at any GitHub repo. Get docs, test plans, executable tests, security findings and a merge-ready Pull Request in minutes.
Workbench is an AI-powered platform for software test analysis and industrialisation. It turns an existing codebase into a complete, traceable, automation-ready test suite, secured against known vulnerabilities and ready to ship as a Pull Request. It combines a fast structured-output AI pipeline with a frontier-model security pass powered by Claude Fable 5, Anthropic's frontier model from the Mythos family for long-horizon cybersecurity reasoning.
Built for teams who want real tests, real security and real traceability — not just pseudo-code.
workbench — analyze
$ workbench analyze https://github.com/acme/payments-svc
→ Cloning repo, detecting toolchain…
✓ Go 1.22 · go test · go modules · godoc
→ Extracting requirements…
✓ 47 requirements (12 high-risk · 18 medium)
→ Generating test cases…
✓ 112 test cases · 94% requirement coverage
→ Building & running test suite…
✓ 108/112 passing · 4 auto-fixed (iter 2/3)
→ Security scan…
✓ 2 CVEs confirmed · 3 suspected · 1 not applicable
→ Fable 5 Discovery…
✓ 3 findings · auth bypass on /admin/* via missing claim check
→ Drafting Pull Request…
✓ PR #482 ready to review
$
A single repository analysis turned into deliverables for development, QA, security and compliance at the same time.
Generate high-quality tests on an existing repository in minutes, not weeks.
Get a complete traceability matrix linking requirements to tests to results.
Surface CVEs and risky code patterns, with AI-validated scope and confirmed exploitability.
One-click export of exportable deliverables (CSV, Markdown, JUnit, audit reports).
Drop in a GitHub URL. Workbench clones and caches the repo, auto-detects the primary language (C/C++, Python, Go, JS/TS, Rust, Java), the project type, test frameworks, package managers, build systems and doc generators.
Generates a structured list of test requirements (R-IDs) from the code, organised by module and classified by risk. You pick test intents across 7 categories: functional, robustness, security, performance, concurrency, regression, compatibility.
Produces concrete test cases (T-IDs) with type, coverage links back to the requirements, evidence citations (file:line) and expected behaviour, preconditions and steps.
Outputs compile-ready tests for the detected framework (GTest, Catch2, pytest, JUnit, Go test, Cargo, Jest, Vitest, Mocha) plus setup/build/test scripts. Syntactic validation and iterative AI auto-correction on failure (up to 3 retries).
Three layers stacked: (1) Pattern-based code analysis (memory issues, command/code injection, weak crypto), (2) Multi-strategy CVE scanning (lockfiles, vendored code, SBOM, imports) with OSV.dev cross-checked against your call graph, and (3) Fable 5 Discovery, a daily security pass powered by Claude Fable 5 (Anthropic's frontier model, Mythos family) reading your code end-to-end to surface logic bugs, auth bypass and taint flows nothing else can see. Each finding ships with a reachability chain from untrusted input to the vulnerable sink. Available on every analyse (3 scans per user per day).
Approve, edit or reject tests inline. One click publishes a merge-ready Pull Request to the source repo with the approved tests, manifests and a CI coverage enforcement script.
Daily allowance of deep security scans powered by Claude Fable 5 (Anthropic's frontier model from the Mythos family, purpose-built for long-horizon cybersecurity reasoning). Reads the codebase end-to-end, surfaces issues other layers can't see, attaches a full reachability chain from untrusted input to the vulnerable sink on every finding. Resilient to safety-classifier declines via automatic fallback.
Vercel/Linear-style read-only URLs for share-without-signup. Owner sees view count and last-used timestamp, revoke at any time.
Connect GitHub once in Settings. Workbench then clones repos your account can see, including private ones in orgs the server token isn't a member of.
Detects test theater by running mull, mutmut or pitest, computing the kill rate and listing surviving mutants per file so you can target real coverage gaps.
Static pattern detection, dependency CVE lookup via OSV.dev with 24h cache, AI anchoring of vulnerable APIs and agentic call-path analysis (confirmed / suspected / not applicable).
Tracks covered vs uncovered requirements, public functions without tests, per-file metrics and a quality score per test (assertions, citations, risk, verification verdict).
Mark a baseline, then detect when cited files change, are deleted or added. Workbench reports impacted requirements so specs never become silently obsolete.
Chat with an agent that greps the code, reads files and artefacts, and answers with file:line citations. Tool calls are streamed live in the UI.
Auto-runs the right doc generator (Doxygen, Sphinx, pdoc, godoc, rustdoc, javadoc, TypeDoc) and hosts the result directly inside the web UI.
One-click export of requirements, test cases, R-ID ↔ T-ID traceability matrix, verification verdicts, JUnit XML results, drift report and a WORKBENCH.md synthesis.
Full async job history with SSE streaming, cooperative cancellation and replay. Prompt caching keeps follow-up analyses at ~10% of the initial cost.
Fable 5 is Anthropic's frontier model from the Mythos family, purpose-built for long-horizon agentic reasoning in cybersecurity. Workbench reserves it for the deep security pass, where it reads the codebase end-to-end and proves each finding with a full reachability chain from untrusted input to the vulnerable sink.
The rest of the pipeline runs on a fast structured-output AI stack tuned for compile-ready output. Fable 5 only enters where its frontier-grade reasoning earns the cost.


KPI tiles for public functions, coverage, CVEs, mutation score and drift. Coverage charts, untested-surface heatmap, security panels, mutation results, test run history, an integrated chat assistant and one-click PR publishing.
$ workbench analyze <url>
$ workbench ask <name>
$ workbench pr <name>
Full pipeline, interactive REPL chat and PR publishing, all from the terminal. Built for CI integration and developer workflows. Use it with public repos out of the box, or pair it with the GitHub Connect flow in the dashboard to scan your private code.
An all-in-one image on GHCR bundles the React UI, FastAPI backend and every toolchain Workbench needs (gcc, cmake, doxygen, JDK, Maven, Go, Rust, Node, gh CLI). Ship it anywhere a container runs.
