Best Free Eval frameworks options.
Candidates with a source-backed free tier, free endpoint, free trial, pretrained download, open weights, open-source license, or self-hosted path. Automated ranking for AI evaluation frameworks, benchmark tooling, observability, guardrails, and quality systems. Rankings use existing source-backed radar data and do not invent prices, benchmark scores, release dates, or capabilities.
Thomeras/agent_detective
Highest automated fit score using source-backed task evidence and radar signals.
Thomeras/agent_detective
Strong candidate when open weights, GitHub, permissive license, or self-hosting signals are visible.
Thomeras/agent_detective
Strong candidate with fresh GitHub, Hub, Product Hunt, or radar momentum.
| # | Candidate | Type | Access | Source | Fit | Updated | Why it matched | Evidence |
|---|---|---|---|---|---|---|---|---|
| 01 | Thomeras/agent_detectivePython repository | repo | Open source | GitHub | Score67 | Jul 29, 2026 | Matched eval framework, llm evaluation, llm eval; 1 source link; access model: Open source; open weights signal | github.com |
| 02 | BlazeUp-AI/ObservalPython repository | repo | Open source | GitHub | Score61 | Jun 9, 2026 | Matched llm evaluation, llm eval, observability; 1 source link; access model: Open source; open weights signal | github.com |
| 03 | Therealdk8890/DProvenanceKitPythonPython repository | repo | Open source | GitHub | Score58 | Jul 15, 2026 | Matched llm evaluation, llm eval, observability; 1 source link; access model: Open source; open weights signal | github.com |
Track Eval frameworks changes
The model, API, pricing and tooling changes that matter, in one morning email with links to the original sources. Free, unsubscribe in one click.
Prefer topic alerts or a custom RSS feed?