Overview
Equilibrium Labs builds the layer that makes AI systems competent at econometrics and honest about economics. Large language models can already write Stata and R. What they lack is applied practice: the judgement about identification, inference, and data identity that separates code which runs from a result which means something.
Everything the project produces is open source. Code is under MIT, written material — including the skills and the method standards — under CC BY 4.0. There is no hosted tier, no paid edition, and no gated component. See governance for why that is a commitment rather than a current state of affairs.
The three problems
1. Econometrics that compiles but identifies nothing
The failure is never syntax. Ask an assistant to estimate the effect of a policy on a panel of firms observed over ten years and the typical response is a pooled OLS regression of the outcome on a treatment dummy and a handful of controls, with the software’s default standard errors, reported as an effect.
Three things have gone wrong at once, and none of them will raise an error. There are no unit fixed effects, so every time-invariant difference between firms — sector, location, management quality — sits in the residual and is correlated with treatment. The standard errors assume observations within a firm are independent across years, which they are not, so the reported precision is overstated, often by a factor of two to four. And the word “effect” has been used for a conditional association that no design supports.
The output looks like an answer. That is precisely the problem: it survives review by anyone who is not already an econometrician.
2. Pre-trained priors presented as consensus
Ask about minimum wages, austerity multipliers, rent control, or trade liberalisation and you get the modal position in the training corpus, delivered with the confidence of arithmetic. A contested literature — one where the empirical estimates genuinely disagree, where the disagreement turns on identification strategy and on which margin is being measured — is flattened into a single sentence beginning “economists generally find”.
The model is not neutral. It is averaged, which is a different thing. Its corpus over-represents Anglo-American academic economics of the past two decades, English-language sources, US institutional arrangements, and published results, which are themselves filtered towards statistical significance. None of that is visible in the answer, and the reader cannot tell a unanimous literature from a split one because the fluency is identical in both cases.
3. Numbers without identity
The subtle failure is not the invented figure. Invented figures are a solved problem in the sense that everyone knows to check them. The failure is the right-looking number from the wrong series.
An assistant compares this quarter’s seasonally adjusted unemployment rate against a figure from a year earlier that was not seasonally adjusted, and reports the difference as a change. Both numbers are real. Both come from the publisher. Both are correctly transcribed. The comparison is meaningless, and it passes every check a reader is likely to apply, because the number cites a source and the source is correct.
The same shape recurs: nominal where real was intended, an index whose base year is not the one assumed, a revised vintage used to evaluate a decision taken when a different figure was on the table, a per-capita figure over the wrong denominator.
What the project ships today
Three things exist in the repository and are usable now.
The skills pack. Agent Skills that install econometric practice into a model’s working context — a directory containing a SKILL.md with YAML frontmatter, installed by copying it into .claude/skills. This is the primary artifact, and it is organised in four groups: practice (the workflow that orders an empirical task, identification, named quasi-experimental designs, panel data, time series, standard errors and inference, regression diagnostics), software (correct modern R and Stata, and reproducibility), reasoning (the moves that make an analysis economic, framework disclosure, and the checklist against the specific ways models get economics wrong), and numbers (series identity and the provenance every figure must carry). The repository’s skills/ directory is the definitive list of what is installable; see installing the skills.
The method standards. Three written documents under method/ in the repository, addressed to people rather than to models: the econometric reporting standard (what an empirical result must contain before it is reportable), the reasoning standard (framework disclosure, contested versus settled questions, positive versus normative), and the provenance rules (the identity a number must carry, and when to refuse rather than report). A skill is an instruction to an agent; a standard is the argument behind the instruction. If you disagree with a skill, argue with the standard.
A validator. python tools/validate_skills.py checks every skill against the format contract — frontmatter keys, name and directory agreement, description length, cross-reference targets, internal link resolution, file length. It is dependency-free and runs in CI on every push to main and on every pull request. A pack that claims to be a standard should be held to one.
That is the complete list. There is no package to install from a registry, no command-line tool other than the validator above, no server, no hosted service, and nothing priced. An MCP server exposing provenance-tracked economic series and a series-identity checker are planned but not built; the roadmap in the repository labels the status of every artifact.
Who it is for
Economists and analysts who already use an AI assistant for empirical work. If you ask a model for R or Stata code, or hand it a dataset with a one-line brief, the pack changes what comes back: the question gets classified before an equation is written, the clustering level is argued from the design rather than taken from the software default, and a descriptive result is not described in causal language.
People building agents that touch economic data. The skills are plain Markdown with a documented frontmatter contract, so they can be loaded by any runtime that reads the Agent Skills format, and read directly by anything that does not. The method standards are the specification if you would rather implement the checks yourself.
Reviewers and teachers, in a smaller way. The standards are a checklist for reading AI-produced empirical work, independent of whether the skills were used to produce it.
What it deliberately does not do
It does not take policy positions. The project has views about method — that the clustering level should match treatment assignment, that staggered adoption breaks two-way fixed effects, that a number without a vintage is not a number. It has no view on whether a minimum wage should be raised, and a skill that starts to acquire one is a bug.
It does not provide data. Nothing here fetches a series, hosts a dataset, or generates synthetic panels. Synthetic data generators were explicitly ruled out: matching marginal moments while getting the dependence structure wrong is worse than no generator at all.
It does not make the model correct. The skills change which questions get asked and which failures get named. They do not verify arithmetic, cannot see your data, and will not catch an error that lies outside the situations they describe. Output still needs review by someone who can review it.
It does not optimise for a confident answer. Several skills instruct the model to decline rather than produce an estimate under ambiguity — an unresolvable unit, a clustering level the data cannot support, an identification claim with no design behind it. An assistant that says “this cannot be identified from these data” is more useful than one that estimates something. If you want an answer regardless, this pack will get in your way, and it is meant to.
Next
Installing the skills has the exact commands for macOS, Linux, and Windows. Your first analysis shows the shape of the difference the pack makes to a single request.
Was this page helpful?
Thanks for the feedback.