What changed
TruePlumb, an open-source project, has released its alpha version (v0.1.0) as a vendor-neutral conformance testbed for AI agent safety controls. The core verification engine, including LTL parser, monitor construction, trace checking with counterexample extraction, and end-of-trace semantics, is implemented and tested. A baseline attack corpus with 26 cases across 9 categories (approval, authorization, prompt injection, memory poisoning, data exfiltration, privilege escalation, resource exhaustion, and termination) is available, along with a corpus schema and validator. The project also includes working CLI tools (verify, atoms, explain), statistics calculation (Wilson intervals, exact McNemar), and a control adapter interface with three reference controls.
Key components still under development include a comprehensive attack corpus, real vendor adapters, and YAML policy file support, along with report generation. The project emphasizes a strict design rule: the scoring path contains zero LLM calls to ensure determinism and prevent gaming. LLMs may be used offline for corpus building but never for verdict decisions. Verification relies on runtime checks using linear temporal logic (LTL) formulas compiled into deterministic finite automata.
Why it matters for builders
This project addresses a critical gap in the AI agent ecosystem: the lack of objective, deterministic methods to validate safety controls. Builders can leverage TruePlumb to move beyond trusting vendor assertions and instead implement rigorous, repeatable testing for their agent's security mechanisms. This empowers developers to build more robust and trustworthy AI agents by providing a clear instrument to measure control effectiveness.
Practical impact
Developers can clone the TruePlumb repository and install it using Python 3.11+. The project offers a CLI for verifying traces against policies, scoring controls over the corpus, and inspecting trace atoms or policy complexity. For instance, trueplumb verify "G(call_refund_api -> F(human_approved))" examples/violating_refund.jsonl can identify violations and provide counterexamples. The score command, such as trueplumb score corpus/ --adapter allowlist, provides detection and false positive rates with confidence intervals, enabling direct comparison between different controls. Builders can also integrate TruePlumb as a library in their Python projects to programmatically check traces or build custom control adapters by subclassing ControlAdapter.
Caveats and source limits
The current release is alpha (v0.1.0), with significant components like the full attack corpus, real vendor adapters, and report generation marked as not started or baseline only. The provided corpus has 26 cases, which is described as a baseline for demonstrating format rather than for competitive ranking. While the verification core is tested and differentially verified against an independent implementation, the broader ecosystem of adapters and comprehensive test cases is still under development. The source does not provide pricing information or specific performance benchmarks against commercial safety solutions.
Sources
Claim check: 5/5 supported claims - 5 evidence links - 100% avg confidence
- TruePlumb is an open-source, vendor-neutral conformance testbed for AI agent safety controls.supported - github.com
- TruePlumb uses judge-free, deterministic verification based on runtime checks of linear temporal logic (LTL) formulas compiled to deterministic finite automata.supported - github.com
- The alpha version includes a working verification core, CLI tools, a baseline attack corpus with 26 cases, and statistics calculation.supported - github.com
- TruePlumb provides metrics such as detection rate and false positive rate with confidence intervals for evaluating safety controls.supported - github.com
- The project is built in the open, with a status table indicating which components are working.supported - github.com
Caveats
- Alpha release.
- Alpha release; corpus is baseline.
- Single-source caution: verify critical details at the linked source.
Radar score 85/100 - how it was calculated
- Reliability 82: GitHub metadata supports source trust
- Freshness 92: Fresh GitHub activity
- Novelty 71: Novelty blends source metadata and enrichment
- Technical 84: Repository technical metadata
- Developer 96: Developer tooling signals
- Ecosystem 66: Developer-oriented GitHub signal
- Confidence 96: Claims have reliable evidence