Equilibrium Labs Docs
v0.1
Changelog GitHub

Press Esc to close

Menu

Questions

The questions a sceptical reader has, answered plainly. Where the honest answer is “not yet measured” or “one person, no guarantee”, that is the answer given.

Is this just a prompt?

Yes, in the sense that instructions supplied to a model are prompts. An Agent Skill is a directory holding a SKILL.md — Markdown with frontmatter — that gets loaded into the model’s context when the task matches its description; there is no fine-tuning, no retrieval index, and no runtime component. Framing that as a dismissal assumes the interesting question is the delivery mechanism. It is not. The interesting question is whether the instructions are correct — whether “cluster at the level of treatment assignment” is the right rule, where its exceptions lie, and whether the estimator recommended for staggered adoption is the current one. That is the part this project spends its effort on, and it is why rules a working econometrician might dispute carry a citation: so a reader can check the claim rather than trust the author.

Does it actually work?

Not measured. The pack’s effect on model behaviour is argued, not demonstrated: the argument is that models fail in specific, repeated ways, that each rule names the failure it prevents, and that a model with the rule in context behaves differently from one without it. That argument is plausible and it is not evidence. What exists today on the verification side is a schema validator (python tools/validate_skills.py) that holds every skill to the frontmatter contract and checks that cross-references resolve — which tests the pack’s consistency, not its effect. A frozen evaluation set of econometric tasks with known-correct answers is listed on the roadmap as considering, and it is described there as the honest gap in the project. Until it exists, treat claims about effectiveness — including the ones on this site — as reasoning rather than measurement.

Why not a Python package?

Because the failure being addressed is judgement, not computation. The computation is solved: statsmodels, linearmodels, fixest, sandwich, reghdfe and ivreghdfe are excellent, actively maintained, and far better than anything this project could write. Nothing here would improve on them and a wrapper around them would add a dependency and subtract nothing. What is missing when a model does applied work is knowing which of those to reach for, what must be true of the data for the answer to mean anything, and what to report so a reader can tell whether it does. None of that is a function call. It is practice, and practice is transmitted in prose.

Will there be a paid version?

No. Everything the project publishes is under MIT for code and CC BY 4.0 for written material, with no component held back — not open-core, not source-available, not free-tier-with-an-edition-behind-it. The reasoning is set out in GOVERNANCE.md and it is a product argument rather than an ideological one: a standard that is partly withheld is not a standard, and guidance that asks to be trusted with published analysis has to be readable, criticisable, and forkable in full by someone who disagrees with it. If the project ever takes funding it will be grants for open work, disclosed rather than converted into restrictions.

Does this work with tools other than Claude Code?

The format is portable Markdown — a directory with a SKILL.md inside it — and it is read today by Claude Code and the Claude Agent SDK, with a growing number of other agent runtimes adopting the same convention. Whether automatic activation works elsewhere depends on that runtime: some load skills by matching the description against the task, some require the skill to be named, and some have no equivalent mechanism at all, in which case the file can still be pasted or referenced directly. Independently of any of that, the content is written to be read by a person. A researcher who reads the method standards and never runs an agent gets most of the value.

Who maintains this?

One person, at low intensity, alongside graduate study. There is no SLA, no response-time guarantee, and no obligation to anything on the roadmap; issues are read but may not be answered quickly. This is stated plainly in GOVERNANCE.md so that nobody builds a workflow around support that does not exist. What is committed is narrower and more durable: nothing already published will be relicensed restrictively, moved behind a paywall, or withdrawn — the existing licences are irrevocable for the versions already released, which means a fork remains possible whatever happens to the maintainer’s attention.

What if a skill is wrong?

Open an issue. This is the single most valuable contribution the project can receive, and it is worth being specific about why: these files are instructions an AI follows literally, so a rule stated too strongly becomes a rule applied where it does not hold, and an exception left unmentioned becomes an exception the model never considers. The highest-value report is therefore not “a skill is missing” but “this skill says X, and X is false when Y” — ideally with a citation, since the project’s stated standard is deference to the applied literature over argument. Concrete transcripts where a model followed a skill and still produced bad econometrics are equally useful, because they show the guidance is present but not operative, which needs a different fix. See CONTRIBUTING.md.

Is this politically neutral?

The project takes no policy positions and has views only about method. That distinction is doing real work, so it is worth spelling out. A view about method is that clustering should match the level of treatment assignment, that staggered adoption breaks two-way fixed effects, that a number without a vintage is not a number. A policy position is a view about whether a minimum wage should be raised, and any skill that starts to acquire one is a bug to be reported. What the reasoning standard asks for is not balance: it treats false balance on a settled question as a defect in exactly the same way it treats false confidence on a contested one. Manufacturing controversy where the applied literature converges makes a model useless on the questions it can actually answer, and presenting a minority framework as suppressed truth is the same error as presenting the majority one as the only one, with the sign flipped.

Do I have to install all of it?

No. The skills are written to work independently and cross-reference each other by name where a handoff matters, so copying the individual directories you want is a supported way to use the pack. The trade-off is that a skill referring to another one that is not installed will simply not get that handoff; nothing breaks, but the chain from design to inference to reporting is looser. Copying everything is the simpler default.

Does the pack need network access, an API key, or a service?

No. There is no package to install, no server, and no hosted component — installation is copying directories into a skills folder, and the files are read from disk by whatever runtime you already use. The one script in the repository, tools/validate_skills.py, checks the pack’s own format; nothing in the skills calls it, and it is standard-library Python if you want to run it yourself. The project deliberately has no operational surface today; a provenance-tracked data server over the Model Context Protocol is on the roadmap as planned and deliberately last, and if it is built it will be self-hosted.

Can I use this in published or commercial work?

Yes. Code is MIT and the written material — the skills and the method standards — is CC BY 4.0, which permits commercial use, modification, and redistribution provided attribution is given. Practically, that means you may adapt a skill to your institution’s conventions and ship the result, as long as the origin is credited. The obligation the licence does not impose, but the project would ask for anyway, is that corrections found along the way come back as issues.

View source on GitHub