Equilibrium Labs Docs
v0.1
Changelog GitHub

Press Esc to close

Menu

Writing a skill

The canonical version of this contract lives in the repository, split between the house style and frontmatter sections of skills/README.md and CONTRIBUTING.md, and it is enforced by tools/validate_skills.py. This page is the expanded copy.

A skill in this pack is a directory containing a SKILL.md with YAML frontmatter, in the Agent Skills format. Optional supporting prose lives in a references/ subdirectory that the model loads only when it needs it. The layout, using the invented survey-weights skill that this page carries as its worked example — it is not part of the pack:

TEXT
skills/
  survey-weights/
    SKILL.md
    references/
      replicate-weights.md

These files are read by a model that will follow them literally, which constrains the prose more than ordinary documentation does: a hedge, a vague sentence or an unverifiable citation behaves differently in a model’s context than on a page a human is skimming. The catalogue of what is already written is in the skill catalogue. Before drafting a new skill, open an issue describing the method and the failure mode an AI shows on it — scope discussion is cheaper than a rewritten draft.

The frontmatter contract

YAML
---
name: survey-weights           # required; lowercase, hyphens, matches the directory
description: What it does...   # required; one unbroken line; third person
license: CC-BY-4.0             # the pack's convention for written material
---

name must be lowercase words separated by single hyphens, at most 64 characters, and identical to the directory name. description must be a single line — a folded multi-line value parses but fails validation, because the model receives it as one activation trigger and a line break in the source has no meaning there. license is checked against CC-BY-4.0 and produces a warning, not an error, if absent or different.

Two further keys are recognised without complaint — allowed-tools and metadata, from the Agent Skills spec. Anything else is reported as a warning, so a typo in a key name does not pass silently.

The description is an activation trigger, not a summary

The description is the only part of a skill the model sees before deciding whether to load it. A summary describes the document; a trigger describes the situations that should pull it into context. Write the second.

That means naming, in one line: what the skill does, and the vocabulary a user would actually type. Include the informal phrasings — “diff-in-diff”, “clustered standard errors”, “is this number real or nominal”, “my regression won’t converge” — not only the textbook terms.

Bad, because nothing in it fires:

YAML
description: A comprehensive guide to best practices for weighting in survey data analysis.

Better, because a request can match it:

YAML
description: Decides whether an estimate should be weighted, which weight to use, and how to compute a standard error under a complex survey design — probability versus analytic weights, when weighting a regression buys consistency and when it only costs precision, and the stratum and primary-sampling-unit variables inference needs. Use when data come from a household or labour force survey, when the user says survey weights, pweight, aweight, svyset, svydesign, strata, PSU or replicate weights, or when a mean or total is reported for a population from a sample that is not self-weighting.

The validator enforces a floor as well as a ceiling: 200 characters minimum, 900 maximum. The floor exists because a description under 200 characters cannot state both what the skill does and when to invoke it. The Agent Skills spec permits up to 1024; the pack is stricter.

Write imperatively, to the model

Instructions, not advice about a topic.

TEXT
Good:  Cluster at the level of treatment assignment.
Bad:   It is important that standard errors be clustered appropriately.

The second sentence is unobjectionable and does nothing. It names no level, so it cannot be executed, and a model that has already chosen the wrong level will read it as confirmation. The first can be followed or violated, which is what makes it a rule.

A rule plus its exception beats a hedge

TEXT
Good:  Use `reghdfe` for multiple fixed effects. Not when you need the
       fixed-effect estimates themselves — then use `areg` or `xtreg, fe`.
Bad:   Various commands may be appropriate depending on context.

Hedging feels safer to write and is worse in use: it leaves the model to pick, which is the behaviour the skill was written to correct. If a rule has boundaries, state them. A rule stated too strongly becomes a rule applied where it does not hold, and an exception left unmentioned becomes an exception the model never considers — so the highest-value contribution to an existing skill is not “this is missing” but “this says X, and X is false when Y”.

Name the failure

Every rule in the pack exists because models fail in a specific, repeated way. Stating the failure is what makes the rule stick, and it lets the model recognise the situation rather than pattern-match the rule.

Mark failures with **Failure:** so they are greppable across the pack:

MARKDOWN
**Failure:** describing fixed effects as "controlling for unobserved
heterogeneity" and stopping. That phrase is true only for *time-invariant*
unobserved heterogeneity. State which confounders the design assumes are
time-invariant, and name at least one plausible time-varying confounder that
survives.

Note the shape: the mistaken behaviour, why it is wrong, and what to do instead. A failure that only says “do not do X” is half a rule. The validator warns when a SKILL.md contains no **Failure:** marker at all.

Refusal is a permitted output, and worth writing explicitly. Where the honest answer is “this cannot be identified from these data” or “these units cannot be reconciled”, say so in the skill and give the model permission to stop. Most skills in the pack carry an explicit refusal section for this reason — thirteen of the fifteen do, under headings such as ## Refusal, ## When to refuse or ## Permission to refuse.

Length is a cost

A skill competes for the model’s context window. Every sentence that is not load-bearing dilutes the ones that are.

The limit is 500 lines for a SKILL.md, enforced as an error. When a topic outgrows it, push detail into references/ and point at the file from the place in SKILL.md where the model would need it — with a sentence saying when to read it, not merely that it exists:

MARKDOWN
For the full catalogue with graphs, including M-bias, Z-bias, the
proxy-confounder case, and when a "bad" control is actually harmless, see
[`references/bad-controls.md`](references/bad-controls.md). Read it whenever
you are deciding a control list of more than two variables.

Reference files have no length limit and no frontmatter. They are ordinary Markdown, loaded on demand. Relative link targets are resolved by the validator against the filesystem, so a pointer to a file that does not exist fails the build rather than silently sending the model nowhere.

Citations

Cite where a rule is contestable. A rule a working econometrician might dispute should carry a reference; a rule from a software manual should name the command and its version behaviour. The project’s standard for contested method is deference to the applied literature, so a change backed by a citation beats a change backed by an argument.

The hard constraint: never attach a citation you cannot vouch for. Not an author-year you believe exists, not a plausible title, not a paper you recall being about roughly this. This project exists because AI systems invent authoritative-sounding references, and a fabricated citation inside a standard that instructs models on honesty is the worst available failure. If you cannot verify it, state the rule without the citation, or write that the point is disputed and name the dispute.

Author-year in prose is the pack’s convention — “Bertrand, Duflo & Mullainathan 2004”, not a footnote apparatus. Package names and functions go in backticks, with the language named where it matters, because vcovHC() in R and , robust in Stata do not default to the same estimator.

Cross-references between skills

Skills are written to work independently and to hand off by name where a handoff matters. Write the handoff as a backticked skill name after a verb the validator recognises — see, load, consult, refer to, hand off to, defer to, delegate to:

MARKDOWN
For everything else — two-way clustering, the choice among candidate levels,
few clusters, randomisation inference — see `standard-errors-and-inference`,
before finalising any inference statement.

Note the see immediately before the backticked name. Phrasing the same handoff as “is in standard-errors-and-inference” reads identically to a human and is invisible to the validator, so a name that later stops resolving would go unnoticed.

References of that form must resolve to a skill directory that exists, or the validator fails. The rule is narrow on purpose: hyphenated lowercase tokens in backticks are also how the skills write R and Stata identifiers, so only the handoff vocabulary is checked.

The catalogue in skills/README.md must list exactly the skills that exist — no entry without a directory, no directory without an entry. Adding a skill means adding its row.

No policy positions

The project has views about method, not about outcomes: that the clustering level should match treatment assignment, that staggered adoption breaks two-way fixed effects, that a number without a vintage is not a number. It has no view about whether a minimum wage should be raised, and any skill that starts to acquire one is a bug.

In practice this means a skill may require the model to name the framework an answer assumes and state what a different school would predict; it may not require the model to prefer one. Where a question is normative dressed as positive, the instruction is to give the positive analysis, name the value judgement, and stop.

Also out of scope: skills for a proprietary tool or dataset most users cannot access, and padding.

Run the validator

BASH
python tools/validate_skills.py            # validate skills/
python tools/validate_skills.py --json     # machine-readable output
python tools/validate_skills.py path/to/skills

It is deliberately dependency-free — standard library only — so there is no environment step. It checks the frontmatter contract, name against the directory, description length and single-line-ness, unrecognised keys, the licence convention, the 500-line limit, the presence of an H1 and of at least one **Failure:** marker, that relative link targets resolve, that cross-referenced skills exist, and that the catalogue and the directory listing agree.

CI runs it on every push to main and on every pull request, along with the validator’s own test suite, so a pull request that fails it will not merge.

A complete minimal skill

The example below is not a skill in the pack. It exists to show the shape: frontmatter that triggers, imperative rules, exceptions stated rather than hedged, named failures, and a refusal condition — inside about seventy lines.

MARKDOWN
---
name: survey-weights
description: Decides whether an estimate should be weighted, which weight to use, and how to compute a standard error under a complex survey design — probability versus analytic weights, when weighting a regression buys consistency and when it only costs precision, and the stratum and primary-sampling-unit variables inference needs. Use when data come from a household or labour force survey, when the user says survey weights, pweight, aweight, svyset, svydesign, strata, PSU or replicate weights, or when a mean or total is reported for a population from a sample that is not self-weighting.
license: CC-BY-4.0
---

# Survey weights

A survey weight is the number of population units an observation stands for.
It is a property of the sampling design, not a modelling option, and using it
correctly is a different decision from using it at all.

## Establish the design before weighting anything

Name four things, from the provider's documentation, before any estimate:

1. The weight variable, and whether it is a probability weight or a
   post-stratification/calibration weight.
2. The stratum variable.
3. The primary sampling unit.
4. Whether replicate weights are supplied instead of strata and PSUs.

If the design is self-weighting — a simple random sample with full response —
say so and skip the rest.

**Failure:** treating the largest numeric column in the file as the weight
because it is called `wt`. Survey files routinely carry person weights,
household weights, and longitudinal weights for the same rows. They answer
different questions and are not interchangeable.

## Weight population quantities; think before weighting a regression

Weight, always, for a descriptive quantity meant to describe the population:
a mean, a total, a share, a quantile, a distribution.

For a regression the decision is not automatic. Weighting is warranted when
sampling probability depends on the outcome, when the estimand is an average
over a population whose composition differs from the sample's, or when effects
are heterogeneous across strata that are sampled at different rates. It is not
warranted merely because a weight exists — under exogenous sampling, weighting
is consistent but less precise than unweighted OLS (Solon, Haider &
Wooldridge 2015).

Report both when they differ materially. A large gap is evidence about
heterogeneity or misspecification, not a menu to choose from.

**Failure:** reporting a weighted and an unweighted coefficient that differ by
a factor of two, choosing one, and not mentioning the other.

## Inference follows the design, not the weight

```r
library(survey)
d <- svydesign(ids = ~psu, strata = ~stratum, weights = ~pweight,
               data = lfs, nest = TRUE)
svymean(~earnings, d)
svyglm(earnings ~ age + educ, design = d)
```

```stata
svyset psu [pweight = pweight], strata(stratum)
svy: mean earnings
svy: regress earnings age educ
```

**Failure:** passing probability weights to `lm(weights = )` or
`feols(weights = )` in R, or using `aweight` in Stata. Those are analytic
weights. They give the right point estimate for a weighted mean and the wrong
standard error for a stratified, clustered design — usually too small, because
the clustering is ignored entirely.

**Failure:** reporting a weighted mean with no standard error at all. A survey
estimate without its design-based uncertainty is a point with no scale.

## Refuse

Say the estimate cannot be produced, and stop, when the strata or PSU
identifiers are absent and the task requires a standard error, or when the
weight's construction cannot be established and the task requires a population
total. Report the unweighted sample quantity, labelled as such, and name the
variable that would resolve it.

Two things to notice. The description names the tools and the words a user would type, so it fires on a request that never says “survey design”. And every rule is either executable or explicitly bounded — there is no sentence whose content is that the topic deserves care.

Before opening a pull request

Contributions are licensed under CC BY 4.0 for written material and MIT for code, matching the rest of the repository. There is no contributor licence agreement and no copyright assignment. The full guidance is in CONTRIBUTING.md.

View source on GitHub