Tools

What is built, and what is not.

Everything here is open source under MIT and CC BY 4.0. There is no hosted edition and no paid tier — see governance for why that is a commitment rather than a stage. The two columns are the maturity split: what you can clone and use today, and what has been committed to but not written. Nothing is described as finished before it is.

Available now

Shipped

In the repository and usable today. Clone it, copy the skills into your agent, run the linter over your analysis, and read the standards that govern both.

  • Skills pack

    Fifteen Agent Skills covering econometric practice, R and Stata, economic reasoning, and number provenance. Install by copying directories.

  • Econometrics linter

    32 rules over R, Stata and Python. Reads a script and reports the wrong clustering level, the silent default, the reflex transformation. No dependencies, no network, no model.

  • Method standards

    The three written standards the skills enforce: econometric reporting, economic reasoning, and provenance. Argue with these, not the skills.

  • Skill validator

    A dependency-free Python checker that holds every skill to the format contract, run in CI on every push.

  • econ-eval

    24 frozen econometric tasks with fully programmatic scoring — no LLM judge. Scores the absence of a wrong reflex, not only the presence of a right answer.

  • MCP server

    Four tools over the same engine, so an agent can lint its own regression before reporting a number. Offline and deterministic; it fetches no data.

  • Harness profile

    Wiring that puts the skills, the linter and the MCP server into whatever agent you already run. Deliberately not an agent — thin harness, fat skills.

  • RL reward adapter

    Turns an econ-eval score into a scalar a training loop can optimise. A rule-engine reward, because text similarity rates a correct answer in the wrong dialect as a failure.

Not built yet

Planned

Committed to but not built — listed here rather than described as finished. The roadmap says why this order, and what was deliberately ruled out.

  • Series-identity CLI

    The record and the comparability check exist as a library and as an MCP tool. There is no eqlabs identity subcommand yet.

  • Worked examples

    Complete analyses from question to write-up, in R and Stata, as plain files any agent can execute through a shell.

  • PyPI release

    The package builds and a Trusted Publishing workflow exists. It has never been run, and the one-time publisher setup is a manual step not yet done.

  • Training a local model

    Gated on the evaluation set discriminating and showing a measurable gap. An honest negative result is a likely outcome and a publishable one.

Read the roadmap