Skip to content

Dependency Skills

A coding agent should know what the libraries a project already depends on can do, and how their authors mean them to be used — for the version the project actually uses. We are trying to get it there in two halves: a library ships a skill inside its own published artifact, and a small lookup on the developer’s machine finds it for the version in use, without drowning the context window.

It is experimental. We have set it up on our own libraries, and the first now carries its skill on Maven Central; the lookup is not published. Everything else on this site is the measurement that shaped it, including the parts that killed our own ideas. How it works is the shortest route to the design.

Three failures, in the order a developer meets them.

Reinvention. An agent writes its own version of something the project already depends on, having never established the library was there.

Drift. A model’s knowledge of a library is averaged across every version it trained on, so it writes against a shape the library no longer has — most stubbornly where something long-established has recently moved or been renamed.

Selection. A project with hundreds of dependencies, a dozen of which could plausibly answer the question, still has to pick one.

A better model does not fix this. The agent reasons correctly over what it can see, and what it can see does not include the library — so this is a distribution problem, not a capability one. And the value of shipping a skill runs inversely to the model’s training exposure: it helps least for the famous library it already half-knows, and most for the small, new, or private one it has never seen.

Reinvention. A service already depends on a resilience library, but the agent adds a hand-rolled retry-with-backoff loop. On the JVM — and worst on Kotlin Multiplatform — nothing told it the library was there: a resolved dependency is an archive in a cache, never unpacked, and none of it is on disk to search.

Drift. A widely-used library moved a core class two releases ago. The model saw the old package ten thousand times in training and the new one once, so it writes the old import — and argues the point when corrected, because its confidence tracks how often a shape appeared, not whether it is still true.

Selection. A Next.js app has moment and date-fns in reach, three HTTP clients, and four argument parsers — all pulled in transitively, all importable. The agent reaches for the one the team does not use, and nothing flags the wrong turn.

The documentation already ships. Across twelve public projects in four ecosystems, the libraries a developer can actually call are most of the dependency graph — and most already publish their own documentation. Two gaps remain. On the JVM, and worst on Kotlin Multiplatform, the library archives everything and never unpacks, so the documentation it ships is not reachable on disk the way an npm or Python package’s is. And everywhere, nothing indexes that documentation against the dependency graph.

Agents read library documentation constantly already — a README, a source file in node_modules, a hover in the IDE, a docs page fetched mid-task. And the whole emerging standards layer is built to do exactly this, deliberately and at scale: the skill formats, the documentation services, the manifest conventions. Every one of them moves prose written by a third party straight into an agent’s context, verbatim — ours included, until we did something about it. That text is written by whoever published the library, and some of it is hostile.

Very little stands in the way of it today. A model’s own resistance to following instructions it finds in library text varies enormously between models of similar capability, within a single vendor’s range — it tracks how a model was trained, not how good it is, so you cannot rely on it. Quoting the text as untrusted data helps, and is defeated outright by putting the same words somewhere else. Of the attacks we ran that actually worked on an agent, the best security scanner caught two of five. And nothing filters at the point the text is read: the same payload written in five languages’ native documentation conventions was delivered perfectly intact by every parser we tried.

So this is not a hazard an index creates. It is one that is already running, on every agent that reads a dependency’s docs — which is all of them.

And on most ecosystems there is no choice about it. An npm package sits exploded in node_modules; a Python one in site-packages. Every README, every JSDoc block, every docstring from every transitive dependency is already in the working tree, and an agent that greps the codebase reads it without anyone deciding that it should. There is no moment where library text is admitted, because it never arrived — it was always there.

The JVM is the odd one out, and for a reason we have spent most of this project complaining about: Gradle and Maven never unpack. A resolved dependency stays an archive in a cache, so none of its prose is on disk to be read incidentally. That is exactly why an index is needed at all — and it is also the only reason there is a door to put a lock on.

What an index changes is that there is finally somewhere to put a control. Library text reaches an agent by a dozen uncontrolled routes, which is much of the reason nothing checks it. An index is the first arrangement where all of it passes through one place — and that is the structural change. The rewriting step and the filter are only what we chose to put there. Nothing in the existing designs is wrong-headed; they simply have nowhere to put a check even if they wanted one.

We are not claiming that generalises cleanly. On an ecosystem that explodes its dependencies onto disk, an index can control what it serves and cannot stop an agent reading node_modules directly. The control is real where the packaging already keeps the text out of the working tree, and partial where it does not — which is an honest limit rather than a feature.

The honest cost is that an index pulls in more third-party text than an agent would have read by itself: hundreds of libraries rather than the handful it happened to open. The trade is breadth for control — every item rewritten before the agent sees it, rather than a few read raw. That trade is worth arguing about, and the measurements behind it are published in full, including where it does not hold.

The lookup we are trying first does less than that, on purpose. It serves only the skills a library’s own authors put in its artifact, only for libraries the project chose, and only for the version it resolved — as written, marked as the authors’ documentation rather than instructions, and never as permission to run or install anything. It does not rewrite them. That narrows what reaches the agent; it does not screen it, and the rewriting step stays with the heavier design.

Text on this site is licensedCC BY 4.0; source code underApache 2.0. © 2026 Brill Pappin.