What changed
StatsPAI has been introduced as a Python-native library focused on causal inference and applied econometrics, aiming to serve as a unified workbench for empirical researchers. It seeks to replicate common workflows found in Stata and R directly within Python, allowing users to load datasets, estimate models, inspect diagnostics, and export results without switching environments. The library offers Stata-style routines such as regress, ivregress, reghdfe, csdid, rdrobust, synth, psmatch2, and table export commands like esttab and outreg2. It also incorporates R-style routines including lm, fixest, did, Synth, DoubleML, MatchIt, modelsummary, and broom.
Key features include adherence to Stata conventions where relevant, such as vce="robust" or vce="cluster firm" for regression models, Stata's small-sample factors, and the ability to use test, lincom, and margins, dydx() post-estimation commands. Outputs are designed to be Python-native, with methods like .summary(), .tidy(), .plot(), and .to_latex(). A significant aspect is its agent-native access; all public functions are registered with machine-readable schemas, and a bundled statspai-mcp server allows integration with MCP clients like Claude Code, Claude Desktop, and Cursor. Companion Stata tooling and skill repositories are also provided to aid in translating and cross-checking workflows.
StatsPAI is available for installation via pip with pip install statspai, supporting Python versions 3.9 to 3.13. Optional extras can be installed for plotting (statspai[plotting]), high-dimensional fixed effects (statspai[fixest]), Bayesian estimation (statspai[bayes]), neural causal models (statspai[neural] or statspai[deepiv]), performance acceleration (statspai[performance]), and spatial analysis (statspai[spatial]). The library includes 14 bundled datasets for offline use, many of which are real published extracts.
Why it matters for builders
StatsPAI's agent-native architecture is a key differentiator, enabling seamless integration with AI agents. This allows developers to build more sophisticated analytical pipelines where AI can directly interact with econometric models and causal inference tools. By providing a unified API that mirrors familiar Stata and R commands, StatsPAI lowers the barrier to entry for Python-first development in econometrics and causal inference, making these powerful statistical techniques more accessible within AI-driven applications.
Practical impact
Developers can leverage StatsPAI to build AI agents capable of performing complex causal inference tasks, such as estimating treatment effects, analyzing panel data, and conducting regression discontinuity designs, all within a Python environment. The library's structured result objects and machine-readable schemas facilitate programmatic access and automation. Builders can explore using sp.regress, sp.ivreg, sp.feols, sp.callaway_santanna, and sp.rdrobust for various estimation needs, and integrate the output with agent frameworks via the MCP server. The ability to translate Stata/R commands using sp.from_stata() and sp.from_r() offers a practical path for migrating existing analyses.
Caveats and source limits
The library's documentation notes that StatsPAI is not a bit-for-bit identical replication of every Stata/R command. The numerical evidence supporting its methods is uneven, with some estimators validated against R/Stata on identical data, others against known-truth simulations, and many only API-stable without explicit numerical parity claims. Each function includes a validation_status field to indicate the level of evidence supporting its numerical output. Users are advised to consult the 'Validation' section before relying on specific numbers for publication. The source does not provide details on performance benchmarks compared to native Stata or R, nor does it specify pricing as it is an open-source project.
Sources
Claim check: 10/10 supported claims - 10 evidence links - 94% avg confidence
- StatsPAI is the first Agent-native Python library for causal inference and applied econometrics.supported - github.com
- StatsPAI offers a unified API, broad cross-method coverage, structured result objects, machine-readable schemas, Skills, an MCP server, and R/Stata parity validation.supported - github.com
- The library supports Stata-style routines like `regress`, `ivregress`, `reghdfe`, `csdid`, `rdrobust`, `synth`, `psmatch2`, and table export commands `esttab`/`outreg2`.supported - github.com
- The library supports R-style routines including `lm`, `fixest`, `did`, `rdrobust`, `Synth`, `DoubleML`, `MatchIt`, `modelsummary`, and `broom`.supported - github.com
- StatsPAI adheres to Stata conventions for variance-covariance matrices (e.g., `vce="robust"`, `vce="cluster firm"`) and post-estimation commands (`test`, `lincom`, `margins, dydx()`).supported - github.com
- Outputs are Python-native with methods like `.summary()`, `.tidy()`, `.plot()`, and `.to_latex()`.supported - github.com
- Every public function is registered with a machine-readable schema, and a bundled `statspai-mcp` server exposes estimators to MCP clients.supported - github.com
- StatsPAI is available for Python 3.9 – 3.13.supported - github.com
- The library bundles 14 datasets for offline use.supported - github.com
- Numerical evidence for some estimators is uneven, with validation status indicating certified/validated evidence versus API-stable breadth.supported - github.com
Caveats
- This claim is based on the project's self-description as the 'first' in its category.
- The extent of 'broad cross-method coverage' and the depth of 'R/Stata parity validation' may vary across different estimators within the library.
- Users should consult the 'Validation' section before relying on specific numbers for publication.
- Single-source caution: verify critical details at the linked source.
Radar score 78/100 - how it was calculated
- Reliability 82: GitHub metadata supports source trust
- Freshness 8: Fresh GitHub release date
- Novelty 78: Novelty blends source metadata and enrichment
- Technical 85: Repository technical metadata
- Developer 96: Developer tooling signals
- Ecosystem 66: Developer-oriented GitHub signal
- Confidence 97: Claims have reliable evidence