Equilibrium Labs Docs
v0.1
Changelog GitHub

Press Esc to close

Menu

Skill catalogue

The canonical catalogue lives in the repository, at skills/README.md. This page is the expanded copy; the repository is what the validator checks.

Fifteen Agent Skills that install applied econometric practice and bias-aware economic reasoning into a model’s working context. Each is a directory containing a SKILL.md with YAML frontmatter, in the Agent Skills format; most also carry a references/ subdirectory holding detail the model loads only when it needs it. Installation is copying directories — see installing the skills.

Each skill declares in its description when it should activate, so the model loads panel-data when it sees a panel and causal-designs when someone claims an effect. You can also invoke one by name. The skills work independently and cross-reference each other where a handoff matters; econometric-workflow is the router that names which sibling to load at each stage of a piece of work.

The pack is v0.1 — usable, versioned and machine-validated, but early. Content will change and skill names may change with it before v1. The SKILL.md in the repository is the source of truth for behaviour; every entry below links to it. To write one of your own, see writing a skill.

Practice — designing and estimating

econometric-workflow

Source

Responsible for. The entry point and router. It orders the work into stages — classify the question as descriptive, predictive or causal; resolve the under-specified request; settle identification; establish the data’s identity; pre-commit the specification; estimate; do inference; run diagnostics; report — and names the sibling skill to load at each. It carries the pack’s specification-search discipline: report the pre-committed specification first, report the whole set of specifications run, never drop a control to obtain significance.

Activates when. Any empirical request arrives — “regress y on x”, “analyse this dataset”, “does X cause Y”, “what drives” — or a panel, survey or time series is handed over with a vague brief.

A failure it prevents. “The user names a regression, so the model treats the specification as given and jumps to lm() / reg. The user named a regression because that is the vocabulary they have, not because they had settled the estimand, the identification, the clustering level, or the sample. Naming a regression is a description of a wish, not a specification.”

identification-strategy

Source

Responsible for. Settling what variation identifies the parameter, what must be true for it to, and what would break it — before any estimating equation is written. It supplies a six-field identification block, a ladder ranking designs from randomisation down to functional-form assumptions, the catalogue of bad controls (post-treatment variables, mediators, colliders, near-instruments), DAGs as control-selection discipline, omitted-variable bias with a signed direction, measurement error, and which estimand a design actually recovers.

Activates when. A user asks “does X cause Y”, “is this causal or just correlation”, “what should I control for”, “is this endogenous”, or proposes an instrument.

A failure it prevents. “The model writes y ~ x + controls, then constructs a story about why the controls make x exogenous. The story is unfalsifiable because it was written to fit the equation.” The companion rule: the kitchen-sink regression is not an identification strategy — it is an admission that there is none.

causal-designs

Source

Responsible for. The assumption set, diagnostic tests, reporting obligations and characteristic fakery of each named quasi-experimental design: difference-in-differences, event studies, staggered adoption, triple differences, instrumental variables, regression discontinuity, synthetic control, and matching or propensity scores. It carries a design-choice table that is read from the variation available to the design that fits, rather than from the method someone wanted to use, followed by the conditions under which no design works and the effect must be declared unidentifiable.

Activates when. Someone says diff-in-diff, parallel trends, event study, IV, 2SLS, first-stage F, exclusion restriction, RD, running variable, bandwidth, synthetic control, donor pool, propensity score or common support — or whenever an analysis claims one thing caused another.

A failure it prevents. Treating a first-stage F above 10 as a licence to report conventional standard errors. The rule of thumb bounds the bias of 2SLS under homoskedasticity in an overidentified model; for a just-identified IV a 5% t-test requires a first-stage F above roughly 104.7 (Lee, McCrary, Moreira & Porter 2022), or an Anderson-Rubin confidence set instead.

panel-data

Source

Responsible for. Fixed-effects practice: what the within transformation absorbs and what it leaves untouched, why fixed effects beat Hausman-driven model choice for causal questions, correlated random effects as the useful middle, incidental parameters in nonlinear models, unbalanced panels and attrition, dynamic panels and Nickell bias, and — at length — the failure of two-way fixed effects under staggered adoption, with the Callaway-Sant’Anna, Sun-Abraham, Borusyak-Jaravel-Spiess and de Chaisemartin-D’Haultfœuille estimators that repair it.

Activates when. The data have repeated observations on the same units over time, or the user says panel, longitudinal, fixed effects, xtreg, reghdfe, feols, staggered rollout or two-way fixed effects.

A failure it prevents. Reporting the two-way fixed effects coefficient as the effect under staggered timing. “The consequence is not ‘some bias’. It is that β can be negative when every single unit-level treatment effect is positive.” A sign flip is possible, so the bias cannot be argued away as probably small.

time-series-econometrics

Source

Responsible for. Stationarity as the precondition for everything: unit-root testing with ADF, KPSS, Phillips-Perron and DF-GLS and their low power against persistent alternatives; spurious regression; differencing versus detrending; cointegration and error-correction models; HAC inference and bandwidth choice; ARIMA and seasonality; VARs, Cholesky ordering and impulse responses; local projections; why Granger causality is predictive precedence rather than causation; structural breaks; and forecast evaluation that is out-of-sample against a mandatory naive benchmark with no look-ahead from revised data.

Activates when. Data are indexed by time, or the user says time series, stationary, unit root, ADF, cointegration, ARIMA, VAR, impulse response, Newey-West, Granger causality or forecast.

A failure it prevents. “Regressing two trending level series, reporting R² = 0.94, and describing it as a strong relationship.” For forecasting, the parallel rule is that a backtest without a benchmark is not evidence — add the random walk, and note that most of the time it wins, and that is the finding.

standard-errors-and-inference

Source

Responsible for. Choosing and defending the standard error rather than accepting the software default: HC0 through HC3 and the cross-package traps, clustering at the level of treatment assignment rather than shopping for a level, few clusters and the wild cluster bootstrap, two-way clustering, Newey-West and Driscoll-Kraay, and serial correlation in DiD panels. It also governs reporting — standard error, interval and p-value together, multiple testing, effect size against a named benchmark, power and minimum detectable effect asked before the result, and why an insignificant coefficient is not evidence of no effect.

Activates when. Any coefficient is about to be reported, or the user says robust or clustered standard errors, vce(cluster), few clusters, wild bootstrap, p-value, multiple testing, Bonferroni, false discovery rate, power calculation or minimum detectable effect.

A failure it prevents. “Reporting vce(cluster state) with 9 states and three stars, with no acknowledgement that the test is invalid. This is the single most common inference failure in applied panel work.” Cluster-robust inference is asymptotic in the number of clusters, so more observations do not help.

regression-diagnostics

Source

Responsible for. Sorting which assumption failure costs unbiasedness, which costs efficiency, and which costs only valid inference — so heteroskedasticity, multicollinearity and non-normal residuals stop being generic alarms. It covers functional form and RESET, logs versus levels with the log-point and retransformation problems, interactions and marginal effects, influence and leverage without a licence to delete observations, VIF folklore, binary and count outcomes, missing data, and winsorising as a declared decision. Its meta-rule: diagnostics run after a pre-committed specification and inform what you report; they are not a licence to search.

Activates when. The user asks about heteroskedasticity, Breusch-Pagan, RESET, log or level, log(y+1), interaction terms, marginal effects, outliers, Cook’s distance, VIF, imputation, winsorising, or whether a specification passes its checks.

A failure it prevents. “Reporting a log-level coefficient of 0.62 as ‘a 62% increase’. It is 86%. This error scales with the effect size, so it is worst exactly where it matters most.” The neighbouring correction: heteroskedasticity does not bias the coefficients, it biases the variance estimator.

Software — writing it up

econometrics-in-r

Source

Responsible for. Correct, modern R: fixest as the default estimator with vcov stated explicitly in every call, sandwich and lmtest when not using fixest, i() and sunab() for event studies, fepois for counts and multiplicative models, panel-aware lag operators, etable and modelsummary for tables, fwildclusterboot for few clusters, and the design packages did, didimputation, HonestDiD, rdrobust and survey. It also enforces project hygiene — here() rather than setwd(), renv, and parallel-safe seeds.

Activates when. The user asks for R code for a regression, or mentions lm, feols, fixest, plm, ivreg, tidyverse, data.table, robust or clustered standard errors in R, an event-study plot, renv or set.seed.

A failure it prevents. Inheriting fixest’s default variance. With at least one fixed effect the default is clustered by the first fixed effect; with none it is i.i.d. — so reordering | year + firm to | firm + year silently changes the standard error of every coefficient. The skill also names the estimation sample as “the most common silent error in R econometrics”: adding a control with 8% missingness estimates that column on a different, non-random sample.

econometrics-in-stata

Source

Responsible for. Correct, modern Stata: reghdfe for multi-way fixed effects and when areg or xtreg, fe is the right choice instead, ivreg2 and ivreghdfe with the weak-identification statistic that matches the variance you asked for, xtset discipline and the time-series operators, csdid, did_imputation and eventstudyinteract for staggered DiD, rdrobust, boottest, factor-variable notation with margins, esttab, and a master do-file that reproduces.

Activates when. The user asks for Stata code or mentions a do-file, regress, reghdfe, xtreg, ivreg2, absorb, vce(cluster), esttab or margins — or asks why two Stata commands give different standard errors.

A failure it prevents. “replace flag = 1 if income > 100000. In Stata missing is larger than any number, so every observation with missing income is flagged. Write if income > 100000 & !missing(income). This is the most common silent data error in Stata and it survives every visual check.” A close second is building age2 and treat_age by hand and then calling margins, which has no way to know those variables move with age and treat.

reproducible-analysis

Source

Responsible for. Structuring a project so every number can be re-derived: immutable data/raw, regenerable data/derived, project-relative paths, seeds at every point randomness enters, environment lockfiles for R (renv), Python (uv, pinned requirements, conda) and Stata (version plus a project-local ado directory), and a per-run manifest recording input hashes, code commit, seeds and package versions. It also covers what to publish when the microdata cannot be shared.

Activates when. A project is being set up or audited, results are meant to be reproducible, a build pipeline is being written, a result cannot be reproduced, or an analysis is about to be declared finished.

A failure it prevents. Any step performed by hand — “editing the raw file. Opening a CSV in Excel to fix a date column, deleting a stray header row, or re-saving it ‘to fix the encoding’.” The corrections are legitimate; they belong in the build script, as code, with a comment saying what was wrong. The governing rule: if the derived directory cannot be deleted and regenerated identically, the pipeline is incomplete.

Reasoning — thinking about economics without inherited priors

economic-reasoning

Source

Responsible for. The moves that make an analysis economic rather than merely statistical. Four questions run before any answer: compared to what, who bears it, what is the budget constraint, does the answer survive everyone doing it. Then opportunity cost, statutory versus economic incidence, partial versus general equilibrium, elasticity plausibility screens, stocks versus flows, levels versus growth rates, average versus marginal, identities that are not behaviour, market clearing as a substantive assumption, mechanism versus effect, the fallacy of composition, self-selection as behaviour, the Lucas critique, and Chesterton’s fence about institutions.

Activates when. Someone asks who bears a tax or a tariff, whether an estimated effect is plausible, what a policy would do, or reads an accounting identity as a causal relationship.

A failure it prevents. “Reporting statutory incidence as economic incidence. ‘The tax is paid by firms’ is a statement about tax collection, not about welfare.” The magnitude screen catches the other common error: a coefficient implying a labour demand elasticity around −4, against a literature clustered near −0.3, is not a large effect — it is almost certainly a broken specification.

paradigm-pluralism

Source

Responsible for. Classifying a question as settled, settled-in-direction, contested or normative, and giving each the treatment it needs rather than the one the others need. It carries a register of what is settled (incidence, comparative advantage, binding price ceilings, externalities, asymmetric information) and what is genuinely open (fiscal multipliers, minimum-wage employment effects, corporate tax incidence, the falling labour share), the requirement that a consensus claim cite a named survey and state whom it sampled, and guardrails against the skill turning the model into a heterodox partisan.

Activates when. The user asks what economists think or whether something is consensus, or raises minimum wages, austerity, rent control, trade, deficits, inflation, inequality or growth.

A failure it prevents. “Applying category C treatment to everything. A model that hedges every claim with ‘different schools of thought disagree’ is useless and is not more honest — it is less, because it obscures the real disagreements by burying them in fake ones.” The mirror failure is both-sidesing a settled question to appear balanced.

bias-audit

Source

Responsible for. A runnable checklist against the specific ways language models get economics wrong, in twelve items — consensus flattening, textbook default, US-centrism and the WEIRD institutional default, recency and salience from the corpus, publication bias and the inflated effect sizes it leaves behind, normative smuggling, efficiency-equity framed as though efficiency were neutral, confidence miscalibration, anchoring on the user’s framing, the just-so story, famous-paper anchoring, and status-quo asymmetry. Each item names the bias, the tell — how it appears in generated text — and a correction that is a specific edit rather than an attitude.

Activates when. Any economic claim, analysis, literature summary or policy discussion is about to be returned.

A failure it prevents. Normative smuggling, which the skill calls the single most common bias in generated economics because the vocabulary carries the judgement: “reforming the labour market” (toward what, judged how), “removing the distortion” (a departure from which benchmark), “the burden of the regulation” (on whom, and what is on the other side of the ledger). The correction is to replace the evaluative term with the mechanism and the distributional consequence, and to name the welfare criterion if one is applied.

Numbers — provenance and identity

series-identity

Source

Responsible for. Pinning down what a number actually is before it is used, compared or deflated. It supplies a nine-slot identity record — series identifier, source, units, adjustment, price basis, frequency, transformation, geography, vintage — and the rules behind each: seasonal adjustment and why an SA figure may never be compared with an NSA one, real versus nominal and which deflator, index rebasing and the non-additivity of chained volumes, per-capita denominators, PPP versus market exchange rates, the three growth-rate conventions, flows versus stocks and gross versus net, methodology breaks (ESA 2010, SNA 2008, NACE revisions, BPM6), vintages and revisions, and how to read provider identifiers rather than guessing them.

Activates when. Any statistic from FRED, Eurostat, the IMF, the OECD, the World Bank or Statistics Denmark is being read, compared, deflated, rebased, spliced or cited — or someone asks “is this real or nominal”, “is this seasonally adjusted”, or “what was GDP at the time”.

A failure it prevents. “The model writes ‘US GDP grew 2.4%’ with no adjustment status, no price basis, no annualisation convention and no vintage. Four different real numbers satisfy that sentence. Two of them differ by more than a percentage point.” Its other signature catch: inventing a plausible series ID, because provider mnemonics are not guessable and a wrong identifier converts an uncertain claim into a false verifiable one.

number-hygiene

Source

Responsible for. The reporting discipline for every figure that leaves an analysis: value, units, series identifier, source, period, vintage, transformation and retrieval date carried as a structural record rather than a URL at the end of a paragraph. It bans stating a checkable statistic from memory, sets significant-figure and rounding rules, requires an interval or a standard error behind every point estimate (or an explicit reason neither can be computed), distinguishes outturn from forecast from projection from scenario, requires “own calculation based on” whenever you derived rather than quoted, and sets the citation order — series ID first, URL last.

Activates when. Numbers are being written up, a table of statistics is being built, data are being cited, or someone asks “where is this from”, “is that right” or “how precise is this”.

A failure it prevents. “The model produces a table of ten country statistics from memory, each individually plausible, several wrong, all formatted identically to retrieved figures.” The rule that follows is absolute: never fabricate a series ID, a table number, a DOI or a URL to support a remembered figure.

Reading the source

Every rule in the pack exists because models fail in a specific, repeated way, and those failures are marked **Failure:** in the Markdown so the whole pack can be searched for them:

BASH
grep -rn '\*\*Failure:\*\*' skills/

That is the fastest way to see what a skill is actually for. The skills are licensed CC BY 4.0; corrections to the econometrics are the most valuable contribution the project can receive, and the authoring rules are in writing a skill.

View source on GitHub