Skip to main content
Back to blog
SkillsAgentic AIDistillationAutomation

Repo-to-Skill: What If Your Code Repositories Became Your AI Agents' Skill Library?

Distill your code repositories into verified agentic skills: the method (BAAI) that improves an AI agent by +134% without touching the model.

9 September 20269 min read
M
Mohamed EL HARCHAOUIAI Engineer Expert

Mohamed is an AI Engineer Expert at Brainum, specialising in agentic system design, RAG pipelines and production AI deployment. He has been helping organisations navigate their AI transformation for over 10 years.

LinkedIn

Your AI agents need skills to do their job well: a SKILL.md file describing the procedure to follow, references to go deeper, scripts to execute. This format has become the standard in a few months. But one question hangs over every deployment: who writes all these skills? When the team hand-writes its tenth skill, fine. By the hundredth — with hundreds of repositories, runbooks and business APIs to turn into reusable operational knowledge — the manual approach collapses.

On September 2, 2026, the Beijing Academy of Artificial Intelligence (BAAI) published "Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills" on arXiv. The paper introduces DisCo, a pipeline that automatically distills code repositories and research papers into verified "skills" — exactly the operational knowledge autonomous agents are missing. The result: the AREX-Skill Library, 5,000+ skills distilled from 1,000 machine learning repositories, organized into 20 domains and 178 capability families.

The number that raises eyebrows in architecture reviews: across MLE-bench's 75 competitions, a GPT-5.5 agent jumps from 31.1% to 72.9% once equipped with these distilled skills — a +134.3% relative gain, without modifying a single model weight. Behind the research result lies a very concrete question for any company industrializing its agents: what if your internal code repositories became your agentic skill library?

1. The blind spot of agent skills: who writes the hundreds of skills your agents need?

In our March article, Skills vs MCP: when to use which?, we explained how to choose between two extension patterns: a Skill encodes a way of reasoning and acting, an MCP server grants access to external systems. That article settled an architecture choice. It left a production question open: who writes these Skills?

Building a skill isn't writing a prompt over coffee: it means extracting operational knowledge from an expert, structuring it into a verifiable procedure, documenting it, maintaining it. A few hours of work per skill, performed by your most expensive people. Yet an agent in production needs skills by the hundreds: one per business domain, one per procedure type, one per integration.

To make the problem tangible: a platform team that wants to equip its deployment agent with the knowledge from 50 internal runbooks quickly runs the numbers — even at two hours per runbook, that's 100 hours of a senior engineer's time before the agent can reliably diagnose a production incident. Multiply that by application repositories, business APIs and compliance procedures, and the manual approach doesn't hold. (This figure is an illustration, not a data point from the BAAI paper.)

At the scale of an internal code estate — application repositories, operations runbooks, business APIs — the question is no longer "how do we write a good skill?" but "who builds our skills, at what cost, with what governance?". That question, never asked until now, is exactly what Repo-To-Skill tackles head-on: after choosing the format, here comes the production chain.

2. Repo-To-Skill (BAAI): distilling a code repository into a verified skill, without human intervention

DisCo, the pipeline presented by BAAI researchers, turns a code repository (or a research paper) into verified skills in four stages:

  • Scoping: restricting the scope of what is worth distilling from the repository — not everything deserves to become a skill.
  • Grounding in evidence: every claim written into a skill must point to evidence in the source — the code, the documentation, the paper's results. No floating knowledge.
  • Skill graph construction: skills are organized into 20 domains and 178 capability families, so they can be found and composed.
  • Verification: every skill is verified before entering the library — a false or outdated skill doesn't get in.

The result is the AREX-Skill Library: 5,000+ skills distilled from 1,000 ML repositories. The GitHub repository VectorSpaceLab/AREX-Skill passed 240 stars within days of publication and appeared in GitHub trending — a sign the community immediately grasped the stakes.

3. Same format as your Claude Skills: SKILL.md, references, scripts — and progressive disclosure built for scale

A detail that changes everything for a company: every skill produced follows a SKILL.md + references + scripts structure. This is structurally the same format we already recommend for your Skills — not a proprietary format to integrate into your systems, but directly compatible with how your agents already load their instructions.

The architectural innovation lies in progressive disclosure. A router first restricts the agent's query to a domain, then a family, then a repository, then a workflow; the agent then loads only the branch it needs. Concretely: facing 5,000 skills, the agent never holds in context more than the handful of files relevant to its current task.

This is the architecture that makes a library of thousands of skills credible without saturating the model's context — the issue that kills most agentic knowledge bases as soon as they grow. A procedure encoded in January is found through its domain and family, not by scrolling a 200-page system prompt.

4. The numbers that matter: +134% on MLE-bench, up to +367% on the hardest tasks

The evaluation covers MLE-bench, a benchmark of 75 machine learning competitions. A GPT-5.5 agent is evaluated without, then with, the AREX-Skill Library:

ScopeWithout distilled skillsWith AREX-SkillRelative gain
75 MLE-bench competitions31.1%72.9%+134.3%
Hardest tasks13.3%62.2%+366.8%

Two takeaways. First, the magnitude: doubling an agent's score — and nearly quintupling it on the hardest tasks — is the kind of jump you'd expect from a model change, not a documentation change. Second, the method: no model weights were modified. The entire gain comes from what the agent knows how to load, and when.

This is the clearest demonstration to date of a thesis we defend at Brainum: operational knowledge is a performance lever distinct from the model itself. You don't have to wait for the next frontier model to transform your agents; you need to give them access to the right knowledge, at the right moment, in a verified format.

5. The real cost: ~$40 per distilled repository, a one-time investment separate from execution cost

Building the library costs about $40 per distilled repository — a one-time investment, separate from the agents' execution cost. Compare that with the expert hours required to hand-write a single skill, and the salaries that come with them.

The math quickly becomes compelling: distilling an estate of 100 internal repositories represents a few thousand dollars of one-time investment — the budget of one week of a senior engineer — whereas hand-writing the same skills would tie up your experts for months. And unlike a classic documentation project, the output is directly usable by your agents, in a verified and organized format.

That cost/result ratio is what makes the topic industrial rather than purely academic: $40 per repository to turn code sitting idle in your GitLab into operational knowledge your agents can mobilize.

6. CIO decision grid: which internal assets (code, runbooks, APIs) deserve to be distilled first

Once the cost/benefit case is made, an operational question remains: where do you start? If the method delivers on its promises, the first decision isn't the AI's — it's yours: which assets to distill first. Four prioritization criteria:

CriterionQuestion to askPriority signal
Frequency of useDoes the task recur daily or weekly?The more frequent the task, the faster the skill pays off
StabilityIs the procedure stable and documented?A stable runbook is an ideal distillation candidate
Cost of errorWhat does an agent failure on this task cost?High cost = reinforced verification, not exclusion
VerifiabilityCan the skill be verified against evidence (tests, logs, docs)?Yes = distillable and governable

Three asset families typically stand out: code repositories (implementation knowledge, conventions, build and deployment procedures), runbooks (operations and diagnostics procedures, often fragmented and dependent on their authors), and business APIs (business rules encapsulated in services nobody dares touch anymore). In all three cases the logic is the same: the knowledge already exists in your estate — what's missing is the chain that turns it into skills your agents can use.

7. How Brainum builds and governs your internal agentic skill library

Choosing the format — Skills or MCP — is an architecture decision you can settle in-house. Building, verifying and governing the production chain at the scale of hundreds, even thousands, of skills is a different job. That's where we come in.

At Brainum, we build and govern skill libraries from your code estate, runbooks and internal APIs:

  • Distillation of your repositories and procedures into skills in the SKILL.md + references + scripts format, compatible with your existing agents.
  • Progressive disclosure to control context cost at scale: routing by domain and family, on-demand loading, a context that doesn't saturate as the library grows.
  • Governance: every skill is grounded in evidence and verified before entering the library; changes are versioned and validated — a skill library is a governed asset, not a folder of prompts.

A question to ask yourself this week: how many of your runbooks, internal repositories and business procedures could become reusable agentic skills today? The answer is counted in days, not months.

Going further

Want to turn your repositories, runbooks and procedures into a governed agentic skill library? Let's discuss your estate.

Share:LinkedIn
Ready

Did this article inspire you?

Let's talk about your AI challenges in a discovery call.

Book a call