Your first analysis
This page is illustrative. It is not a transcript of a recorded session, and it contains no coefficients, no datasets, and no results, because the point is the shape of the difference rather than any particular number. Model output varies between runs; what follows is the pattern to expect, not a script.
The request, one sentence, of the kind that arrives constantly:
Check whether the minimum wage increase reduced employment using this panel.
Assume the panel is what such a panel usually is: units observed over time — regions, or firms, or establishments — with a wage floor that rose in different places at different dates, and an employment measure taken from a published statistical series.
Without the pack
An unassisted assistant typically does something close to this, and does it immediately.
It reads the request as a specification. “Whether the minimum wage increase reduced employment” names an outcome, a treatment, and a direction, so the model treats the estimating equation as given and writes it:
m <- lm(employment ~ min_wage_increase + controls, data = panel)
summary(m)It reports the coefficient, its default standard error, and a p-value. If it is being careful it adds year dummies. The prose around the estimate says the increase “reduced employment by” some amount, or “had no significant effect on” employment.
Four things have gone wrong before the first line of output.
The question was never classified. Nobody established whether this is a descriptive, predictive, or causal question. The request is causal — it asks what employment would have been otherwise — but nothing in the analysis is doing the work that a causal claim requires.
There is no design. The comparison being made is between units with higher and lower minimum wages, conditional on controls. That comparison is confounded by everything that determines where wage floors rise: local labour-market tightness, political composition, cost of living, prior employment trends. Adding controls is not an identification strategy, because it names no assumption that could be false.
The panel is not being used as a panel. With no unit fixed effects, every time-invariant difference between regions sits in the error term and is correlated with treatment. Even with unit and time fixed effects, staggered adoption dates make the two-way fixed effects coefficient a weighted average of comparisons that includes already-treated units serving as controls.
The standard error is the software default. It assumes observations are independent, when policy varies at the level of the region and employment shocks persist within a region for years. The reported precision is overstated, frequently by a factor of two to four.
The output is fluent, formatted, and wrong in a way that is invisible unless you already know what to look for. That asymmetry — fluent output produced faster than it can be reviewed — is the whole reason this project exists.
With the pack installed
The same sentence produces a different first move: the model does not write the regression.
The question gets classified, and estimation is blocked
econometric-workflow activates on requests of exactly this shape — a vague brief with a named outcome, or a dataset handed over with one line of instruction. Its stop rule is explicit: do not write the regression, classify the question first.
The question is causal. That classification is not cosmetic — it determines what the answer owes the reader. A causal question requires a source of variation, an identifying assumption, falsification tests, and a named estimand. A descriptive question requires none of those, and demanding them would be a category error. The workflow skill states which one it is answering in a sentence, then hands off.
It also surfaces the small number of things that cannot be guessed: what the unit of observation is, what the unit of policy variation is, which parameter is wanted — an average effect, an effect on the treated, an effect at a margin — and which population the answer is meant to describe. It asks at most a few of these and proceeds under stated assumptions for the rest, rather than turning the request into an interrogation.
The identification strategy has to exist before the equation
identification-strategy loads because the question is causal. It requires, in prose, before any estimating equation:
- The source of variation. What varies, for whom, and why. “Some regions raised their wage floor and others did not, at dates determined by state legislation” is a source of variation. “Minimum wage differs across regions” is a restatement of the data.
- The counterfactual. What the treated units are being compared to, and why that comparison stands in for what would have happened to them.
- The identifying assumption, stated as a claim about unobservables that could be false. Parallel trends is such a claim. “We control for relevant covariates” is not.
- The threats, and what was done about each — including reverse causality, since wage floors are raised in response to labour-market conditions, and the sign that bias would take.
- The estimand recovered. ATT is not ATE, and the design decides which one is available.
This is also where the honest answer can be that no credible identification is available from these data. economic-reasoning reinforces the same discipline from the other direction: it forces the counterfactual to be stated explicitly, asks who actually bears the cost of the policy as distinct from who is nominally subject to it, and applies an order-of-magnitude plausibility check to whatever estimate eventually emerges.
The design’s assumptions come with it
If the answer to the identification step is a difference-in-differences comparison, causal-designs supplies that design’s specific obligations rather than generic advice: what parallel trends means and what evidence bears on it, why pre-treatment coefficients that are individually insignificant are a statement about power rather than evidence of parallel trends, and what an event-study plot must show.
The panel skill flags the staggered-adoption problem
panel-data activates on repeated observations of the same units. With treatment turning on at different dates in different places, it raises the failure that models reproduce most often: under staggered adoption the two-way fixed effects estimand is a weighted average of all possible two-by-two comparisons in the data, and some of those use already-treated units as controls for later-treated ones. Some weights are negative. The consequence is not a small bias — the coefficient can take the opposite sign to every single unit-level treatment effect.
So the pack does two things. It diagnoses first, decomposing the estimate to report how much of it comes from forbidden comparisons and what share of the weights are negative. Then, if that share is non-trivial, it moves to an estimator built for heterogeneous timing — Callaway and Sant’Anna, Sun and Abraham, Borusyak, Jaravel and Spiess, or de Chaisemartin and D’Haultfœuille — each of which targets cohort-by-period effects and aggregates with weights you choose rather than weights the data impose. In R that is did::att_gt() with aggte(), or fixest::sunab() inside feols. It also requires the event-study path to be shown rather than a single number, because effects varying is the entire premise of switching estimators.
econometrics-in-r supplies the idioms that keep this from going silently wrong — the estimation sample staying constant across specifications, the never-treated coding convention differing between packages, the cluster argument not being forgotten.
The clustering level is argued, not defaulted
standard-errors-and-inference applies a rule: cluster at the level at which treatment is assigned, or at the level at which the sample was drawn, whichever is coarser. Not the level of the outcome. Not the unit of observation. Not the level that produces the largest standard error.
Here the policy varies by region and year, so the clustering level is the region — even though the outcomes are establishments or individuals, and even though clustering more finely would give a comfortable number of clusters. Clustering more finely than treatment assignment because the correct level is inconvenient is named as a failure mode, not a compromise.
That raises the next problem, and the skill raises it unprompted: if the number of regions is small, cluster-robust inference does not work. It is asymptotic in the number of clusters, not the number of observations, and below roughly thirty to fifty clusters a nominal five per cent test can reject far more often than that. A million establishment-years in twelve regions gives twelve effective observations for inference. The response is wild cluster bootstrap or randomisation inference, not a larger sample.
The same skill governs how the result is reported: the standard error, the confidence interval, and the p-value together; economic magnitude distinguished from statistical significance; and an insignificant coefficient not reported as evidence of no effect, since that requires a power calculation and a minimum detectable effect that the design may not support.
The employment series has to have an identity
series-identity activates because a published statistic is entering the analysis. Before the employment measure is used, it establishes what the number actually is: seasonally adjusted or not, and the same on both sides of every comparison; the level of aggregation and whether a methodology break falls inside the sample window; whether the series was revised and which vintage is loaded; whether a per-capita or per-establishment denominator is involved and what it is. If any of that cannot be resolved, refusing is a permitted output — the rules behind this are in the provenance rules.
This is the least dramatic of the checks and one of the most consequential. A methodology break or a change in seasonal adjustment inside the sample can move the series by more than the effect anyone is trying to estimate.
number-hygiene then governs how every figure appears in the write-up: value, units, series identifier, source, period, transformation, and vintage carried as a structural record rather than a formatting preference; no statistic quoted from memory when it is checkable; no point estimate without an interval or an explicit reason one cannot be computed; and rounding that does not quietly change the conclusion.
The write-up follows the reporting standard
What comes back is not a coefficient in a sentence. It follows the econometric reporting standard: the question classified as causal; the identification strategy and its assumption stated before the equation; the data with source, identifier, units, period, and vintage; the sample construction as a sequence of steps with counts; the specification, the estimator, and why that estimator rather than two-way fixed effects; the estimate with its standard error and interval and the clustering level named; the pre-period path and the falsification evidence; and the threats that remain unaddressed, said plainly.
What actually changed
Nothing in the list above is exotic. It is what a careful applied economist does, and most of it appears in any good empirical seminar. The difference is that an unassisted model skips it, at speed, in fluent prose, and the skipping is invisible in the output.
Three specific changes are worth naming.
The order changed. Identification is settled before estimation rather than defended after it. Several requirements in the reporting standard are worthless if produced in the wrong order — a specification pre-committed after seeing the result is not pre-committed.
Refusal became available. The pack gives the model explicit permission to say that a parameter cannot be identified from these data, that units cannot be reconciled, or that the clustering level the design requires leaves too few clusters for reliable inference. Some requests should end there, and an assistant that estimates something regardless is worse than one that stops.
The claim matched the evidence. Where the design does not license a causal reading, the language becomes associational — “employment differs by” rather than “the increase reduced employment”. That single substitution is what separates a reportable descriptive result from a misreported causal one.
Trying it
Install the pack, then hand an assistant a real request from your own work and watch which skills fire. The clearest signal that the pack is working is that the first response contains no code.
If a skill you expected did not activate, name it directly — see how activation works. If a rule in a skill is wrong, or has a well-known exception it does not mention, open an issue on the repository. Corrections to the econometrics are the most valuable contribution the project can receive.
Was this page helpful?
Thanks for the feedback.