Econometric reporting standard
The canonical version of this standard lives in the repository, at method/econometric-reporting-standard.md. This page is the rendered copy.
What an empirical result must contain before it is reportable.
This is deliberately close to what a good applied journal already requires. The contribution is not novelty; it is that the requirements are written down in a form an agent can be held to, and that the order is enforced — several of the items below are worthless if produced after the result rather than before it.
1. The question, classified
State whether the question is descriptive, predictive, or causal. The three have different standards and are routinely conflated.
- A descriptive question asks what is true in the data. It needs representative data and honest uncertainty. It does not need an identification strategy, and demanding one is a category error.
- A predictive question asks what value to expect. It needs out-of-sample evaluation against a benchmark. It does not need unbiased coefficients — a biased model can predict well, and interpreting its coefficients causally is the error.
- A causal question asks what would happen under an intervention. It needs everything in section 2, and without it the result is descriptive whatever the language around it says.
A result that answers a descriptive question in causal language is misreported even when every number in it is correct.
2. The identification strategy, stated before estimation
For causal questions only:
- The source of variation. What varies, for whom, and why — in words, before any equation.
- The counterfactual. What the treated units are being compared to, and why that comparison stands in for what would have happened.
- The identifying assumption, stated as a claim about unobservables that could be false. “Parallel trends” is such a claim. “We control for relevant covariates” is not — it names no assumption that could fail.
- The threats, and what was done about each. Where nothing could be done, say so.
- The parameter recovered. ATE, ATT, LATE, or an effect local to a cutoff. These are different quantities and the design determines which one you get.
3. The data, with identity
Every variable used carries source, identifier, units, period, and — for anything revised — vintage. The full rules are in the provenance rules.
Additionally:
- Sample construction, as a sequence of steps with the count remaining at each. Every exclusion is a decision that must be visible.
- Missingness: how much, where, and what was done. Listwise deletion is a choice with assumptions, not a default.
- The estimation sample must be constant across specifications that are presented as comparable. Silently dropping observations when a control with missing values is added produces a table whose columns are not comparable — report N per column so this is visible.
4. The specification, pre-committed
State the specification before seeing its result. Then report it, whatever it shows.
- Robustness means the same question under different reasonable choices, with all of them reported. It does not mean estimating many specifications and reporting the ones that survive.
- If the specification changed after seeing results, say so and say why. This is sometimes legitimate — a coding error, a discovered data problem — and it is only illegitimate when concealed.
- Controls require justification. A control that is a consequence of treatment, or a common effect of treatment and outcome, introduces bias rather than removing it. “It was available” is not a reason.
5. Inference matched to the design
- Standard errors that match how the data were generated. Cluster at the level at which treatment was assigned or sampling was done, and state the level and the number of clusters. Default homoskedastic errors are almost never right in applied work.
- With few clusters, asymptotics fail. Below roughly 30–50, cluster-robust inference over-rejects substantially; use wild cluster bootstrap or randomisation inference, and say which.
- Report the confidence interval, not only the point estimate and stars.
- Report the effect size in interpretable units against a benchmark — a share of the mean, a share of a standard deviation, a comparison to a known estimate. “Statistically significant” describes the standard error, not the finding.
- A null is a finding with content. Report what the interval rules out. “No significant effect” from an underpowered design and “precisely estimated zero” are opposite results and must not be written the same way.
- Multiple outcomes require an adjustment or a pre-specified primary outcome. Say which.
6. Diagnostics that inform, not diagnostics that select
Run diagnostics after the pre-committed specification. Report what they show. They may change how the result is described; they do not license a search for a better one.
Deleting influential observations, transforming until a test passes, or dropping a control because significance appears are all specification search regardless of the diagnostic that motivated them. If an observation is removed, report the result with and without it.
7. Limitations that could change the conclusion
Not a ritual paragraph. Each limitation states what it would take to overturn the result:
- The assumption most likely to fail, and the direction the estimate would move if it did.
- The population the estimate applies to, and where it should not be extrapolated.
- What data would settle the question if it existed.
8. Reproducibility
Code, seed, package versions, and the exact path from raw data to reported table. Where the data cannot be shared, publish the code, the manifest, and the derived aggregates — an unshareable dataset is not an excuse for an unshareable method.
The refusal clause
If the data cannot identify the parameter the question asks about, the correct output is to say so and stop. Producing an estimate with a caveat attached is not equivalent: the estimate gets quoted and the caveat does not.
The three situations that require refusal rather than a hedged estimate:
- No credible source of variation. The design does not exist and no robustness check can create one.
- Irreconcilable units or definitions. The variables do not measure comparable quantities and no transformation makes them comparable.
- Inference impossible at the required level. Too few clusters for any valid method, or a design whose standard errors cannot be computed.
Related
- Reasoning standard — framework disclosure, contested versus settled questions, and the positive/normative line.
- Provenance rules — the identity a number must carry, and when to refuse rather than report.
Licensed CC BY 4.0.
Was this page helpful?
Thanks for the feedback.