Equilibrium Labs Docs
v0.1
Changelog GitHub

Press Esc to close

Menu

Econometric reporting standard

The canonical version of this standard lives in the repository, at method/econometric-reporting-standard.md. This page is the rendered copy.

What an empirical result must contain before it is reportable.

This is deliberately close to what a good applied journal already requires. The contribution is not novelty; it is that the requirements are written down in a form an agent can be held to, and that the order is enforced — several of the items below are worthless if produced after the result rather than before it.

1. The question, classified

State whether the question is descriptive, predictive, or causal. The three have different standards and are routinely conflated.

A result that answers a descriptive question in causal language is misreported even when every number in it is correct.

2. The identification strategy, stated before estimation

For causal questions only:

3. The data, with identity

Every variable used carries source, identifier, units, period, and — for anything revised — vintage. The full rules are in the provenance rules.

Additionally:

4. The specification, pre-committed

State the specification before seeing its result. Then report it, whatever it shows.

5. Inference matched to the design

6. Diagnostics that inform, not diagnostics that select

Run diagnostics after the pre-committed specification. Report what they show. They may change how the result is described; they do not license a search for a better one.

Deleting influential observations, transforming until a test passes, or dropping a control because significance appears are all specification search regardless of the diagnostic that motivated them. If an observation is removed, report the result with and without it.

7. Limitations that could change the conclusion

Not a ritual paragraph. Each limitation states what it would take to overturn the result:

8. Reproducibility

Code, seed, package versions, and the exact path from raw data to reported table. Where the data cannot be shared, publish the code, the manifest, and the derived aggregates — an unshareable dataset is not an excuse for an unshareable method.

The refusal clause

If the data cannot identify the parameter the question asks about, the correct output is to say so and stop. Producing an estimate with a caveat attached is not equivalent: the estimate gets quoted and the caveat does not.

The three situations that require refusal rather than a hedged estimate:

  1. No credible source of variation. The design does not exist and no robustness check can create one.
  2. Irreconcilable units or definitions. The variables do not measure comparable quantities and no transformation makes them comparable.
  3. Inference impossible at the required level. Too few clusters for any valid method, or a design whose standard errors cannot be computed.

Licensed CC BY 4.0.

View source on GitHub