Equilibrium Labs Docs
v0.1
Changelog GitHub

Press Esc to close

Menu

Glossary

The vocabulary a reader of these docs actually meets, in alphabetical order. Definitions are written for someone comfortable with data who is not necessarily an econometrician. Where a skill in the pack is responsible for a term, its source is linked.

A

Agent Skill. A directory containing a SKILL.md file — YAML frontmatter with a name and a description, followed by Markdown instructions — that an agent runtime loads into the model’s context when the description matches what the user is doing. Optional references/ files sit alongside it and are loaded only when that level of detail is needed. Installing one means copying the directory; nothing is compiled and there is no runtime dependency. The format is documented at code.claude.com/docs/en/skills.

ATE, ATT and LATE. Three different average treatment effects, routinely confused. The ATE (average treatment effect) is the average across the whole population; the ATT (average treatment effect on the treated) is the average among units that actually received treatment; the LATE (local average treatment effect) is the average among compliers, the units whose treatment status an instrument moved. Designs identify different ones — difference-in-differences typically an ATT, instrumental variables a LATE, regression discontinuity an effect at the cutoff — so reporting one and discussing another is a misstatement rather than a simplification.

B

Bad control. A variable that introduces bias when it is included rather than removing it — typically a variable that is itself affected by the treatment, or a common consequence of both treatment and outcome (a collider). Controlling for the first blocks part of the effect being estimated; controlling for the second opens a spurious association. That a variable was available in the dataset is not a justification for including it.

C

Chain-linking. The method national accounts use to build a volume (“real”) series: growth between adjacent periods is computed at the earlier period’s prices, and those growth rates are multiplied into an index. It avoids the distortion of holding one distant base year’s prices fixed, at the cost of additivity — chain-linked components do not sum to the chain-linked total, which is a property of the method rather than an error in the data. See series-identity.

Clustering (of standard errors). Allowing for correlation between observations within a group — a school, a firm, a state — rather than assuming every observation is independent. The rule is to cluster at the level at which treatment was assigned or sampling was carried out, and to report that level together with the number of clusters. Clustering too finely understates uncertainty; clustering too coarsely can leave too few clusters for the asymptotics that cluster-robust inference relies on. See standard-errors-and-inference.

Cointegration. Two or more non-stationary series can be individually trending yet move together, so that some linear combination of them is stationary. When that holds the series are cointegrated, the long-run relationship is real, and the system has an error-correction representation. When it does not, regressing one integrated series on another produces high R-squared values and significant coefficients that mean nothing — the classic spurious regression.

Contested, settled, framework-dependent. The three-way classification the reasoning standard requires before an economic question is answered. A settled question is answered directly, with no manufactured balance; a contested one is answered while saying what the answer depends on; a framework-dependent one is answered while naming the framework assumed and what a different one would predict. Most questions are settled, and treating them otherwise trades a bias for a uselessness.

D

Difference-in-differences (DiD). A design that compares the change in an outcome for a treated group against the change for a comparison group over the same period, on the assumption that the two would otherwise have moved in parallel. The estimate is the difference of those two differences. Naming the method is not a design: the comparison group, the timing, and the counterfactual claim have to be stated before anything is estimated. See causal-designs.

E

Estimand. The quantity you are trying to estimate, stated in words — whose effect, of what, on what, over what horizon — before any estimator is chosen. Distinguishing the estimand from the estimator matters because two regressions with identical syntax can target different quantities depending on the variation they use.

Event study. A difference-in-differences in which the single before/after indicator is replaced by one coefficient per period relative to treatment, plotted as a path. It is the primary evidence for a DiD result and the primary place that result gets read dishonestly: one period is always omitted as the baseline, the endpoints are estimated from fewer units than the middle, and a flat pre-period with wide intervals is uninformative rather than reassuring.

F

Fixed effects. Absorbing everything constant within a unit (or within a period) by giving each one its own intercept, so the estimate is identified only by variation within units over time. Mechanically this is the within transformation: each unit’s own mean is subtracted from every variable, which removes time-invariant differences and amplifies the effect of classical measurement error. Fixed effects do not “control for unobserved heterogeneity” in general — anything time-varying survives untouched, and anything that does not vary within a unit cannot be estimated at all. See panel-data.

H

HAC standard errors. Heteroskedasticity- and autocorrelation-consistent standard errors — Newey-West for a single series, Driscoll-Kraay for panels — which allow the error term to be correlated across nearby periods. They are the time-series counterpart to clustering, and they matter because serial correlation in an outcome inflates apparent significance badly.

I

Identification. The argument that the variation being used in the data recovers the parameter the question is about, and not something else. It is a claim about the world, not a property of a regression command: what varies, for whom, why, what it is being compared against, and what would have to be true for that comparison to stand in for the unobserved counterfactual. An identifying assumption must be capable of being false — “we control for relevant covariates” names no assumption that could fail. See identification-strategy.

Instrumental variable (IV). A variable that shifts the treatment of interest, is as good as randomly assigned, and affects the outcome only through that treatment. The third condition — the exclusion restriction — cannot be tested; it has to be argued channel by channel from institutional knowledge. Where treatment effects differ across units, IV recovers a LATE for compliers, not the population average effect.

P

Parallel trends. The identifying assumption behind difference-in-differences: in the absence of treatment, the treated group’s average outcome would have followed the same path as the comparison group’s. This is a statement about a counterfactual that is never observed. It is not the same claim as “pre-treatment trends looked parallel”, and a pre-trend test that fails to reject does not establish it — particularly since such tests are usually underpowered.

Percentage point (pp). The unit for a difference between two percentages. A move from 4 per cent to 6 per cent is a rise of 2 percentage points, or of 50 per cent; calling it “a 2 per cent rise” is wrong and changes conclusions. Use percentage points for differences between percentages, always. See the provenance rules.

Provenance. The record of where a number came from and what was done to it: source, series identifier, units including scale and currency, observation period, transformation, vintage, and retrieval date. Provenance is treated here as a structural requirement rather than a citation style, because the errors it prevents are invisible in the output — a sentence reads identically whether or not the figure in it was deflated, adjusted, or of the right vintage.

R

Real and nominal. A nominal figure is measured in the prices of its own period; a real figure has been adjusted for price change and is expressed in the prices of a reference period. “Constant prices” names a base year, not a price index, and different deflators give materially different answers over a decade or more. A real series therefore is not fully identified until the deflator and the reference period are both stated.

Refusal. Declining to produce an estimate or report a figure, rather than producing one with a caveat attached. The standards treat refusal as a valid output in specific situations — no credible source of variation, units that cannot be reconciled, too few clusters for any valid inference, an undeterminable adjustment status, an unknown deflator for a real series, a figure recalled rather than retrieved. The reasoning is that an estimate travels and its hedge does not. See the econometric reporting standard and the provenance rules.

Regression discontinuity (RD). A design in which treatment is assigned by whether a continuous score crosses a known threshold, so that units just either side of the cutoff are comparable. It identifies the effect at the cutoff, for units at the cutoff, and nothing else — not the average effect of the programme, and not the effect of moving the threshold any material distance. A fuzzy RD, where crossing only shifts the probability of treatment, is an instrumental-variables design and carries every IV obligation.

S

Seasonal adjustment. The removal of estimated seasonal and calendar components from a series so that adjacent periods can be compared. Whether a series is adjusted cannot be inferred from its values, and an adjusted figure must never be compared against an unadjusted one — the apparent discrepancy is the seasonal factor. Adjustment is itself revised as new observations arrive, so the adjusted history changes even when the raw data do not.

Series identifier. The publisher’s own code for a series — a FRED series ID, a Eurostat dataset and dimension combination, an IMF indicator code. It is treated as the primary key for a number because URLs rot and titles are ambiguous, while identifiers persist and resolve to exactly one series with one adjustment status, one price basis, and one frequency.

Specification search. Trying specifications until one produces an acceptable result, then reporting that one. It includes the versions that feel principled at the time: deleting influential observations, transforming a variable until a diagnostic passes, adding unit-specific trends to rescue a failing pre-trend. The inference reported afterwards does not account for the search, which is why the specification is pre-committed and then reported whatever it shows. See regression-diagnostics.

Staggered adoption. Treatment that turns on at different dates for different units, rather than all at once. It is the common case in policy evaluation and it is the case in which conventional two-way fixed effects breaks, because the estimator uses already-treated units as controls for newly-treated ones. Heterogeneity-robust estimators exist for it and should be used rather than assumed unnecessary.

Stationarity. A series is stationary when its mean, variance, and autocovariances do not depend on when you look. Many macroeconomic series are not: they contain a unit root and wander without reverting to a fixed level. Regressing one such series on another unrelated one reliably produces a large t-statistic, which is why the order of integration is established before a time-series regression is interpreted.

Synthetic control. A design for one treated unit (or a few) with many untreated donors and a long pre-treatment period: a weighted average of donors is built to reproduce the treated unit’s pre-treatment path, and the post-treatment gap between the two is the estimate. Weights are non-negative and sum to one, so the synthetic unit cannot extrapolate beyond the donor pool. Conventional standard errors do not apply — inference comes from placebo permutations across donors and across dates.

T

Two-way fixed effects (TWFE). A regression with both unit and period fixed effects — the standard workhorse for panel data. With a single treatment date it is a well-behaved difference-in-differences estimator. Under staggered adoption with treatment effects that vary over time, it becomes a weighted average of all possible two-group, two-period comparisons in which some weights are negative, so the reported coefficient can carry the opposite sign to every underlying unit-level effect.

V

Vintage. The particular release a figure came from. Macroeconomic data is revised, sometimes by enough to change the sign of a quarter, so a statement about what was known at a point in time requires the figure that was published at that time rather than today’s. Vintage is one of the two fields that most often gets dropped — the other being transformation — and its absence is invisible in the resulting sentence.

W

Wild cluster bootstrap. A resampling procedure used when the number of clusters is small — below roughly 30 to 50 — and cluster-robust standard errors over-reject, sometimes badly. It repeatedly reweights residuals at the cluster level under the null hypothesis to build a reference distribution for the test statistic. Randomisation inference is the alternative where treatment assignment was itself random or as-if random.

View source on GitHub