Workbench

    Point Workbench at any GitHub repo. Get docs, test plans, executable tests, security findings and a merge-ready Pull Request in minutes.

    Workbench is an AI-powered platform for software test analysis and industrialisation. It turns an existing codebase into a complete, traceable, automation-ready test suite, secured against known vulnerabilities and ready to ship as a Pull Request. It combines a fast structured-output AI pipeline with a frontier-model security pass powered by Claude Fable 5, Anthropic's frontier model from the Mythos family for long-horizon cybersecurity reasoning.

    From a GitHub URL to a merge-ready PR in minutes
    Real executable tests, auto-corrected by AI
    Code security analysis + dependency CVE scan
    Fable 5 Discovery: Anthropic frontier model deep security scan with reachability chains
    Public share links: read-only report URLs, no account required
    Full traceability: requirements ↔ tests ↔ code
    Web UI for everyone, CLI for developers

    Built for teams who want real tests, real security and real traceability — not just pseudo-code.

    workbench — analyze

    $ workbench analyze https://github.com/acme/payments-svc

    → Cloning repo, detecting toolchain…

    ✓ Go 1.22 · go test · go modules · godoc

    → Extracting requirements…

    ✓ 47 requirements (12 high-risk · 18 medium)

    → Generating test cases…

    ✓ 112 test cases · 94% requirement coverage

    → Building & running test suite…

    ✓ 108/112 passing · 4 auto-fixed (iter 2/3)

    → Security scan…

    ✓ 2 CVEs confirmed · 3 suspected · 1 not applicable

    → Fable 5 Discovery…

    ✓ 3 findings · auth bypass on /admin/* via missing claim check

    → Drafting Pull Request…

    ✓ PR #482 ready to review

    $

    One workflow, four teams.

    A single repository analysis turned into deliverables for development, QA, security and compliance at the same time.

    01

    Development teams

    Generate high-quality tests on an existing repository in minutes, not weeks.

    02

    QA teams

    Get a complete traceability matrix linking requirements to tests to results.

    03

    Security teams

    Surface CVEs and risky code patterns, with AI-validated scope and confirmed exploitability.

    04

    Compliance & audit

    One-click export of exportable deliverables (CSV, Markdown, JUnit, audit reports).

    From URL to PR, step by step.

    1. 1

      Point at a GitHub repo

      Drop in a GitHub URL. Workbench clones and caches the repo, auto-detects the primary language (C/C++, Python, Go, JS/TS, Rust, Java), the project type, test frameworks, package managers, build systems and doc generators.

    2. 2

      Extract requirements

      Generates a structured list of test requirements (R-IDs) from the code, organised by module and classified by risk. You pick test intents across 7 categories: functional, robustness, security, performance, concurrency, regression, compatibility.

    3. 3

      Generate test cases

      Produces concrete test cases (T-IDs) with type, coverage links back to the requirements, evidence citations (file:line) and expected behaviour, preconditions and steps.

    4. 4

      Generate executable code

      Outputs compile-ready tests for the detected framework (GTest, Catch2, pytest, JUnit, Go test, Cargo, Jest, Vitest, Mocha) plus setup/build/test scripts. Syntactic validation and iterative AI auto-correction on failure (up to 3 retries).

    5. 5

      Security & CVE scan

      Three layers stacked: (1) Pattern-based code analysis (memory issues, command/code injection, weak crypto), (2) Multi-strategy CVE scanning (lockfiles, vendored code, SBOM, imports) with OSV.dev cross-checked against your call graph, and (3) Fable 5 Discovery, a daily security pass powered by Claude Fable 5 (Anthropic's frontier model, Mythos family) reading your code end-to-end to surface logic bugs, auth bypass and taint flows nothing else can see. Each finding ships with a reachability chain from untrusted input to the vulnerable sink. Available on every analyse (3 scans per user per day).

    6. 6

      Review and publish a PR

      Approve, edit or reject tests inline. One click publishes a merge-ready Pull Request to the source repo with the approved tests, manifests and a CI coverage enforcement script.

    Everything you need to industrialise testing.

    Fable 5 Discovery

    Daily allowance of deep security scans powered by Claude Fable 5 (Anthropic's frontier model from the Mythos family, purpose-built for long-horizon cybersecurity reasoning). Reads the codebase end-to-end, surfaces issues other layers can't see, attaches a full reachability chain from untrusted input to the vulnerable sink on every finding. Resilient to safety-classifier declines via automatic fallback.

    Public share links

    Vercel/Linear-style read-only URLs for share-without-signup. Owner sees view count and last-used timestamp, revoke at any time.

    Private repo support

    Connect GitHub once in Settings. Workbench then clones repos your account can see, including private ones in orgs the server token isn't a member of.

    Mutation testing

    Detects test theater by running mull, mutmut or pitest, computing the kill rate and listing surviving mutants per file so you can target real coverage gaps.

    Multi-layer security

    Static pattern detection, dependency CVE lookup via OSV.dev with 24h cache, AI anchoring of vulnerable APIs and agentic call-path analysis (confirmed / suspected / not applicable).

    Coverage & quality scoring

    Tracks covered vs uncovered requirements, public functions without tests, per-file metrics and a quality score per test (assertions, citations, risk, verification verdict).

    Drift detection

    Mark a baseline, then detect when cited files change, are deleted or added. Workbench reports impacted requirements so specs never become silently obsolete.

    Repo-grounded AI assistant

    Chat with an agent that greps the code, reads files and artefacts, and answers with file:line citations. Tool calls are streamed live in the UI.

    Documentation generation

    Auto-runs the right doc generator (Doxygen, Sphinx, pdoc, godoc, rustdoc, javadoc, TypeDoc) and hosts the result directly inside the web UI.

    Compliance pack

    One-click export of requirements, test cases, R-ID ↔ T-ID traceability matrix, verification verdicts, JUnit XML results, drift report and a WORKBENCH.md synthesis.

    Run history & cost control

    Full async job history with SSE streaming, cooperative cancellation and replay. Prompt caching keeps follow-up analyses at ~10% of the initial cost.

    Powered by Anthropic

    Claude Fable 5, Mythos-class.

    Fable 5 is Anthropic's frontier model from the Mythos family, purpose-built for long-horizon agentic reasoning in cybersecurity. Workbench reserves it for the deep security pass, where it reads the codebase end-to-end and proves each finding with a full reachability chain from untrusted input to the vulnerable sink.

    The rest of the pipeline runs on a fast structured-output AI stack tuned for compile-ready output. Fable 5 only enters where its frontier-grade reasoning earns the cost.

    Claude Fable 5 and Mythos 5 announcement artwork by Anthropic: the number five composed of butterflies
    Artwork: Anthropic — Claude Fable 5 & Mythos 5 announcement, June 2026.

    Web UI for everyone. CLI for developers.

    Workbench Summary view: run status, Fable 5 Discovery card, KPI tiles (requirements, test cases, coverage, security findings), languages, test outcomes and security & CVE charts

    Web dashboard

    KPI tiles for public functions, coverage, CVEs, mutation score and drift. Coverage charts, untested-surface heatmap, security panels, mutation results, test run history, an integrated chat assistant and one-click PR publishing.

    • Requirements Atlas with risk/gap filters
    • Inline test editing and bulk approval
    • Live job streaming with SSE and replay
    • Live prompts counter: pre-run estimate plus running total during the analyse, so customers see their quota burn in real time

    Command line

    $ workbench analyze <url>

    $ workbench ask <name>

    $ workbench pr <name>

    Full pipeline, interactive REPL chat and PR publishing, all from the terminal. Built for CI integration and developer workflows. Use it with public repos out of the box, or pair it with the GitHub Connect flow in the dashboard to scan your private code.

    What makes it different.

    01From URL to PR in minutes, a true end-to-end workflow
    02Tests that actually compile and run, fixed by the AI when they break
    03Three layers of security: code patterns + CVE + agentic scope validation
    04Fable 5 Discovery built in: Anthropic's Mythos-family frontier model, Claude Fable 5, runs the deep security pass, bringing frontier-grade reachability proof to your code review, not a generic SAST. Other Workbench phases use a fast structured-output AI pipeline tuned for compile-ready output; Fable 5 is reserved for the security layer where its long-horizon agentic reasoning earns the cost.
    05Mutation testing built in to measure real test quality, not just coverage
    06Full traceability: requirements ↔ tests ↔ code ↔ results, audit-ready
    07Drift detection keeps requirements aligned with the codebase over time
    08AI assistant grounded in your code with file:line citations
    09Cost control via prompt-cache reuse (~10% of initial cost on reruns)
    10Dual surface: web UI for non-technical users, CLI for developers

    One Docker image. Any cloud.

    An all-in-one image on GHCR bundles the React UI, FastAPI backend and every toolchain Workbench needs (gcc, cmake, doxygen, JDK, Maven, Go, Rust, Node, gh CLI). Ship it anywhere a container runs.

    Docker ComposeAWS FargateAWS App RunnerGoogle Cloud RunAzure Container AppsGHCR

    See Workbench on your own repo.

    Book a demo and we will point Workbench at a repository of your choice, public or under NDA, and walk you through every artefact it produces.

    Chat with AI Assistant