Full document← Back to the map

The Quantium Skill Standard

A proposal for how skills are authored, published, discovered and maintained across Quantium

Status: Draft for discussion Owner: Adam Witanowski, Executive Manager — AI Solutions Aligns to: Strategy Execution Plan, Enabler C.1 (Technical Foundations — shared infrastructure, Marketplace, Q.Eval, CI/CD)

Suggested working name: Dewey, after the librarian who solved classification.


In one page

Problem. ~1,700 skills exist across laptops and repos with no curation. Five document skills, six presentation skills, none canonical. Users can't tell what's right, fixes don't propagate, nothing is owned.

Key finding. Distribution, versioning, auto-update, group targeting, drift prevention and per-skill telemetry all already exist as product capability. We should not build any of it. What's missing is curation, a quality bar, ownership, a contribution path for non-technical staff, and a defect loop.

Proposal. One repository, four levels — org, team, community, personal — running through one set of machinery, with evidence-driven escalation between them. Exactly one skill per capability at org level. Federated ownership: the publisher owns, others report bugs. Quality enforced by CI: structural lint, capability collision check, and behavioural evaluation via Q.Eval.

Interaction vision. Users never see GitHub. Everything they touch is either Claude or Slack. The tooling is itself four skills, distributed by the marketplace it governs.

Day one. A new consultant opens their laptop and the right skills are already there. There is one presentation skill, so "am I using the right one" never arises.

Effort. Phases 0–2 are configuration and process, twelve weeks, no new infrastructure, no dependency on vendor roadmap. The only real software build is a Q.Eval adapter, if one is needed.


1. The problem

Skills have spread faster than any governance around them.

WhereRoughly how manyGoverned?
Individual laptops and personal claude.ai accounts~1,500No
One central GitHub repo~30Partially
Three engineering repos for Claude Code~200Per-repo, inconsistent

The count is not the problem. The duplication is. Most of these are forks of each other, each carrying one small improvement that never made it back. The cost lands in three places:

At current growth this becomes unmanageable within two quarters. The window to standardise cheaply is now, while the volume is still mostly noise.

2. What we are not building

The distribution layer already exists as product capability.

Anthropic's organisation plugin marketplaces distribute curated skill bundles to everyone, or to specific groups, from a private GitHub repository that syncs automatically. Distributed skills appear in claude.ai chat, in Claude Desktop and in Cowork; Claude Code users get the same repo via one command. Each bundle can be set to Required, Installed by default, Available for install, or Not available, and Enterprise group targeting overrides that per team. Users cannot edit organisation-managed skills, which eliminates silent forking at the source. Per-skill telemetry exists too (§11).

That is distribution, versioning, auto-update, group scoping, drift prevention and measurement — roughly 70% of what we would otherwise have specified and built.

What is genuinely missing: curation, a quality bar, an ownership model, a contribution path for non-technical staff, and a defect loop. These are process problems with a small amount of glue around them.

We reviewed the emerging commercial category — AWS Agent Registry, Google's Agent Registry, JFrog's Agent Skills Registry, TrueFoundry. Recommendation: do not buy one yet. They are built for multi-vendor agent and MCP sprawl, which we do not have. Revisit when skill provenance and signing become an AIMS compliance requirement, or when we are genuinely running three or more agent vendors in production.

3. Surfaces: where skills are consumed

A common question is why we plan for chat users at all rather than putting everyone on Cowork.

It is not a licensing decision. On Enterprise, the seat fee covers access to Claude on web, desktop and mobile plus Claude Code and Cowork; usage is not included in the seat fee and every token consumed in chat, Claude Code or Cowork is billed at standard API rates on top. Cowork is not a separate purchase.

It is a consumption decision. Agentic sessions do far more work per task — screenshots, sub-agent coordination, multi-step file operations — with reported multiples of five to twenty times a comparable chat interaction. For someone whose day is questions and drafting, that spend buys nothing.

It is mostly a fit decision. Cowork earns its cost when there is a multi-step task with real artefacts to delegate. Adoption should follow demonstrated need.

Two consequences:

  1. Cowork must be enabled at the organisation level regardless, because it is a prerequisite for plugin marketplaces to function at all. Enablement is not optional; adoption can be uneven.
  2. The same bundles serve all three surfaces. Group targeting configured once carries across chat, Cowork and Claude Code.

4. Personas

PersonaSurfaceWhat they need
Priya — Consultantclaude.ai chat only. Non-technical. No GitHub account and shouldn't have one.The right skills already present. To know they're official. To flag when one is wrong.
Marcus — EngineerClaude Code, some Cowork. Comfortable with git.Repo-scoped skills, fast local iteration, no ceremony for routine updates.
Deb — Skill ownerChat. May be a Finance lead who wrote the board-pack skill, not an engineer.To know her skill is used, to hear about defects, to fix it without learning git.
Raj — Team leadChat and Cowork.Skills that encode his team's way of working, distributed to his team automatically.
Sam — Curator (AI Lab)Claude Code.To arbitrate contested slots on evidence, not politics. To not become the bottleneck.
Alex — Platform adminAdmin console.Sane defaults, group mapping from the IdP, an audit trail, no per-person configuration.

The constraint that follows from Priya and Deb: GitHub is a database and an audit trail, never a user interface. If a non-technical person ever has to open github.com, that is a design defect, not a training gap.

5. The target experience

Day one

Priya opens her laptop for the first time. The skills her role needs are already installed and enabled — she took no action, made no choices, and read no documentation. There is exactly one presentation skill, one document skill, one client-proposal skill. The question "am I using the right one" does not occur to her, because there is only one.

Marcus runs one command to register the marketplace and gets the engineering bundle. He keeps his own local skills for his own work; they simply aren't in the official set.

Time to productive: zero for Priya, under five minutes for Marcus.

The moment a skill is wrong

Priya notices the deck skill used last year's client colour palette. She says so, in chat. Claude captures the skill version and what went wrong, files it, and posts it to #skills-feedback tagged to Deb. Deb replies in Slack. Neither opens a browser tab they didn't already have open.

The moment someone has something worth sharing

Deb has been getting good results from a prompt she keeps re-pasting. She asks Claude to turn it into a skill. Claude interviews her — what is it for, when should it fire, show me an output you liked — writes it to the house standard, checks it against the taxonomy and tells her the truth: "we already have an official board-pack skill; yours does one thing ours doesn't. Want me to propose that as a change to the official one instead?"

Most contributions should end in a redirect rather than a new skill. That conversation is the highest-leverage thing in this design, and only an agent can have it at scale.

The moment something outgrows its level

Raj's team has been using their client-reporting skill for two months. Two other teams have independently adopted it. The monthly sweep flags it, Q.Eval runs it against the reporting slot, and a curator confirms. It becomes an org standard. Raj is notified; he remains the owner. Nobody applied for anything.

6. Use cases

  1. Get set up — new starter receives their role's skill set with no action.
  2. Use the right skill — one skill per capability at org level, so selection is unambiguous.
  3. Author a skill — conversational, guided to the standard, no markdown or YAML knowledge required.
  4. Donate an existing skill — migrate what's on laptops into the governed set.
  5. Publish to the community library — share something useful without it having to be a firm standard.
  6. Encode a team's way of working — a skill that is canonical for one team and not the firm.
  7. Escalate on evidence — promotion driven by usage and evaluation, not lobbying.
  8. Report a defect — in chat, routed to the owner, without a ticketing system.
  9. Retire gracefully — demotion rather than deletion when a skill stops earning its slot.

7. The model

Four levels, one machinery

OrgTeamCommunityPersonal
MeansHow Quantium does thisHow this team does thisSomeone found this usefulMine
Locationorg/ in repoteam/<name>/ in repocommunity/ in repoStays with the user
DistributionRequired, org-wideRequired, that IdP groupAvailable for installNot distributed
Slot ruleExactly one per slotSub-slots onlySub-slots onlyn/a
Quality gateLint + collision + full Q.EvalLint + collision + evalLint + minimal evalAuthoring-time checks
OwnershipNamed owner, review SLATeam ownerNamed owner, best effortn/a
SupportFirm-supportedTeam-supportedAs-isNone

Team level exists because a team's client-reporting skill is genuinely canonical for that team and genuinely not a firm standard. Placing it in community understates it; placing it in org is wrong. It also turns escalation into a ladder rather than a jump, which is what makes automatic promotion tractable — each step is a small evidentiary leap.

Mechanically, team level costs nothing new: it is Required-for-that-IdP-group, same marketplace, same sync, different distribution preference.

The community library is a first-class destination, not a holding pen. Most useful skills should never become firm standards. Treating them as failed candidates guarantees they go back onto laptops.

Personal is registered, not stored

Personal skills should not live in the repository. Three reasons:

The resolution is storage versus registration. Personal skills stay where they live. The skill-author and donate-my-skill skills run the same lint and safety checks at authoring time, client-side, before the skill is ever used, and register a manifest entry — name, description, owner, hash, proposed capability slot — without content. We get the safety check and the inventory signal (three people wrote the same thing → that slot needs an official answer) without the repo bloat or the exposure.

Note also that on Enterprise, third-party skills and plugins users upload or edit are already scanned for malicious content by the platform. Our checks add prompt-injection, credential-leakage and tool-permission review on top; we are not starting from unprotected.

Consequently, escalation out of personal is user-initiated. Automatic escalation starts at community, because we cannot promote what we should not be able to see.

The escalation ladder

From → ToTriggerDecided by
Personal → CommunityUser donatesUser
Community → TeamUsage concentrated in one groupTeam owner accepts
Community → OrgUsage spread across ≥3 groups; slot free or wins on evalQ.Eval + curator
Team → OrgIndependently adopted by ≥3 teamsQ.Eval + curator
Org → CommunityUsage falling, or eval regressionAutomatic, with notice
Any → ArchivedNo use in two quartersAutomatic, with notice

Three things about this table matter more than the thresholds:

Thresholds trigger evaluation, not promotion. Usage measures popularity, not correctness. A widely used skill producing subtly wrong output is worse than an unused one. The sweep nominates, Q.Eval decides, a curator confirms. Total human effort: one meeting a month.

Demotion is the row nobody builds and everyone needs. An org skill that stops earning its slot drops a level rather than disappearing. This frees the slot for a better candidate and removes the political cost of retiring something.

Team → Org via independent adoption is the strongest signal in the system. Three teams arriving at the same skill without coordinating is better evidence than any usage count.

Alongside usage, the monthly sweep watches two other signals: convergence (several community skills clustering on one capability — the clearest sign a slot needs an official answer) and absence (heavy community activity in a slot the org library doesn't cover). These make the lower levels a demand signal for what the standard is missing.

Ownership

Per-folder ownership via CODEOWNERS at every level. The publishing team or individual owns their skill: they review changes and fix defects, and ownership survives promotion. Ownership is not exclusivity — anyone can propose a change to anyone's skill and the owner reviews it. This gives us an internal open-source dynamic rather than a central team that becomes a queue.

8. The skills that run the system

The tooling is itself made of skills, distributed by the same marketplace it governs. There is no web application to build, log into or maintain — which is also why the non-technical experience works.

skill-author — "help me make a skill"

Distribution: Required, everyone.

The refusal case is the valuable one: "we have this; yours differs in one respect; propose that as a change instead." Most authoring sessions should end here.

donate-my-skill — "I already have one of these"

Distribution: Required, everyone. Two behaviours by surface.

The two paths produce very different input quality. Tag provenance on intake so the queue can be triaged.

report-skill-issue — "this skill got it wrong"

Distribution: Required, everyone.

skill-telemetry — for curators

Distribution: Available, Analytics role holders only.

9. System architecture

  AUTHORING                    GOVERNANCE                      DISTRIBUTION
  ─────────                    ──────────                      ────────────

  Priya  (chat)  ─┐
  Deb    (chat)  ─┼─► skill-author ─────┐
  Raj    (chat)  ─┤   donate-my-skill   │
  Marcus (Code)  ─┘   report-skill-issue│
        │                               ▼
        │                   #skills-intake  (Slack)
        │                               │
        │             curator or agent opens PR
        │                               │
   ┌────▼─────────────┐                 ▼
   │ PERSONAL         │   ┌─────────────────────────────────────┐
   │ stays local /    │   │  quantium-skills   (private repo)   │
   │ user account     │   │  ├── .claude-plugin/                │
   │                  │   │  │      marketplace.json            │
   │ manifest entry   ├──►│  ├── org/         <plugin folders>  │
   │ only: name,      │   │  ├── team/<name>/ <plugin folders>  │
   │ owner, hash,     │   │  ├── community/   <plugin folders>  │
   │ proposed slot    │   │  ├── taxonomy.yml                   │
   └──────────────────┘   │  ├── registry.yml  (personal index) │
                          │  └── CODEOWNERS                     │
                          └──────────────┬──────────────────────┘
                                         │  pull request
                          ┌──────────────▼──────────────────────┐
                          │  CI gates  (blocking)               │
                          │   1. structural lint                │
                          │   2. collision + sub-slot check     │
                          │   3. Q.Eval behavioural suite       │
                          └──────────────┬──────────────────────┘
                                         │  merge + version bump
                          ┌──────────────▼──────────────────────┐
                          │  Org plugin marketplace             │
                          │  (auto-sync on merge)               │
                          └──────────────┬──────────────────────┘
                                         │  group mapping from IdP / SCIM
          ┌──────────────────────────────┼──────────────────────────────┐
          ▼                              ▼                              ▼
     claude.ai chat                   Cowork                      Claude Code
     Desktop chat tab

Measurement, feedback and escalation:

   Analytics API /skills ─┐
                          ├─► owner digest job ─► Slack DM per CODEOWNER   (monthly)
   open issues per folder ┘

   report-skill-issue ──► GitHub issue ──► #skills-feedback, owner tagged  (immediate)

   monthly sweep ──► usage + convergence + absence + independent adoption
                          │
                          ├─► promotion candidates ─► Q.Eval ─► curator ─► PR
                          └─► demotion candidates  ─► notice ─► automatic PR

Everything in the middle is invisible to Priya, Deb and Raj. They interact with two surfaces: Claude, and Slack.

Component inventory:

ComponentStatusEffort
Org plugin marketplace, group targeting, auto-syncExistsConfiguration
Per-skill usage and cost analyticsExistsConfiguration
Platform malicious-content scanningExistsNone
Marketplace repo, taxonomy, CODEOWNERSBuildSmall
Structural linterBuild1–2 days
Collision and sub-slot checkBuildSmall
Q.Eval integration for skill artefactsConfirm, then buildUnknown — §10
The four skills (§8)BuildSkills, not software
Slack intake and feedback botBuildSmall
Owner digest job and monthly sweepBuildSmall — scheduled script

10. The quality gate

Every change entering the repo passes three gates.

Gate 1 — Structural lint (mechanical, seconds). Valid frontmatter; name kebab-case within 64 characters; description within the 200-character limit; CODEOWNERS entry present; version bumped; folder structure correct. A required CI check. This matters more than it sounds, because an increasing share of skills are written by agents, and structural checks that feel pedantic to a human are exactly what neither the contributing agent nor the reviewing human reliably catches.

Gate 2 — Collision and sub-slot check (automated). Two rules:

The second rule is what makes four levels safe. Without it, an org skill and a team skill both describing themselves as "for making presentations" turn selection into a coin flip — rebuilding the original problem inside a single user's session. It is a naming convention backed by CI, and it is the difference between four levels working and four levels being worse than two.

Gate 3 — Behavioural evaluation (Q.Eval). Two measures:

Depth scales with level: authoring-time checks only for personal, three to five cases for community, a fuller suite for team, a full suite plus regression on every change for org.

Two secondary benefits worth naming to the business:

Open item: I believe the internal tool here is Q.Eval, listed under Enabler C.1. This needs confirming, along with whether it supports skill artefacts as an evaluation target or needs an adapter. If the latter, that adapter is the only meaningful software build in this proposal.

11. Telemetry and the owner loop

Per-skill telemetry is not something we need to build. The Analytics view already reports users, cost per use and total uses for every skill in the organisation, filterable by group and by product surface, with CSV export; the Enterprise Analytics API exposes the same through a /skills endpoint alongside plugin adoption. Members can see their own usage by product, model and skill in their settings.

The gap is routing, not collection. Analytics access is scoped to owners and a custom Analytics role and is organisation-wide rather than per-team, so we cannot hand a skill owner an API key.

The mechanism: a scheduled job pulls the /skills endpoint, joins the results against CODEOWNERS and open issues, and pushes each owner their own slice as a Slack message.

You own board-pack-builder (team level, Finance). Last month: 47 people, 210 runs, $0.60 per run. Two other teams have adopted it. Two open issues. Last updated five months ago. It's a candidate for org promotion — want me to run the eval?

Owners get exactly their data. Nobody gets analytics access they shouldn't have. The same pull feeds the monthly sweep, the escalation ladder and the staleness model — one job serving four purposes.

Two constraints to plan around. Analytics data is aggregated per organisation per day with a one-to-two day lag and a 90-day retention window, so warehouse an extract from Phase 2 if we want longer trend history. And per-skill analytics is an Enterprise capability — on Team it is dashboards and CSV only, with no API.

12. Roadmap

PhaseTimingOutcomeSuccess measure
0 — Foundations2 weeksTenancy prerequisites confirmed and enabled. Capability taxonomy and sub-slot naming convention written (one page). Harvest run: filesystem sweep for engineering skills, "which three would you hate to lose?" survey for everyone else.Taxonomy agreed. Candidate list under 100.
1 — Day one worksWeeks 3–6quantium-skills repo live. The twelve launch slots filled (Appendix A). Required bundles mapped to IdP groups. Structural lint and collision check in CI.New starter productive with zero setup. Duplicate skills in active use trending down.
2 — Contribution and feedbackWeeks 7–12The four skills shipped. Community and team levels open. Personal registration live. Slack intake and feedback channels running. Owner digest job running. Telemetry warehousing started.First 20 community contributions. First 3 team-level skills. First defect resolved without anyone opening GitHub.
3 — Quality gate and escalationQ1Q.Eval wired into CI as a blocking gate. Monthly sweep operating across all escalation paths, including demotion.Every org skill has an eval suite. First promotion and first demotion both driven by evidence.
4 — Self-sustainingOngoingStaleness model live. Agent proactively offers to skill-ify repeated manual work. Archival running.Skills improve without central intervention. Retirement rate roughly matches creation rate.

Phases 0–2 need no new infrastructure and no dependency on the vendor roadmap. Telemetry lands in Phase 2 rather than later because the underlying data already exists.

13. Asks and dependencies

Internal

From Anthropic (via our ANZ contact) — desirable, not blocking:

  1. An admin inventory view of personal skills — names, descriptions, owner, without content. Usage is already covered; discovery of what exists is not, and it is what our personal-registration workaround exists to substitute for.
  2. Org-level skill management API for CI/CD. Already requested publicly by others; adding our weight is cheap.
  3. Self-serve export of personal skills, so donation doesn't depend on people still having the original zip.

14. Risks

RiskMitigation
Org and team skills compete for the same trigger, rebuilding the problem inside a single session.Sub-slot rule enforced in CI (Gate 2). This is the primary risk of the four-level model.
Four levels are more ceremony than the firm will tolerate.Levels differ only by folder and distribution preference. One linter, one eval harness, one telemetry join. Complexity is in the rules, not the machinery.
Two distribution paths (directory sharing and the marketplace) both become "official".Directory sharing maps to community. The marketplace is authoritative. State it explicitly and repeatedly.
We import the mess. A harvest optimised for completeness gives a landfill with a search box.Harvest by nomination, not by sweep. The 1,400 we don't capture dying quietly is the success metric.
Promotion by usage promotes popular skills rather than correct ones.Threshold triggers evaluation, not promotion. Q.Eval decides.
Community level is perceived as a failure tier, so people don't publish there.Give it its own identity and browsable surface. Never describe it as "not yet official".
Central team becomes the bottleneck.Automate gatekeeping via lint and eval. Humans review additions and escalations only, not routine updates.
Turning off user-created skills too early kills the grassroots supply.Keep creation on. Govern the org level, not the sandbox.
Telemetry beyond 90 days is lost.Warehouse the extract from Phase 2.

15. Decisions requested

  1. Endorse the approach: curate and govern, do not build distribution infrastructure.
  2. Approve the four-level model with personal registered rather than stored, and evidence-driven escalation including demotion.
  3. Confirm Q.Eval as the quality gate and fund any adapter work in Phase 3.
  4. Approve federated (publisher) ownership rather than centralised ownership.
  5. Nominate the taxonomy authors — the one piece that is genuinely intellectual work and cannot be delegated to tooling.

Appendix A — The proposed org library

The test for org level: would it be wrong for a team to do this differently? If a team could legitimately differ, it belongs at team level. This is what keeps the org library at thirty rather than three hundred.

SlotWhat it doesLaunch
Orientation
im-new-hereFirst-30-days orientation: what exists, house norms, where to go
quantium-glossaryInternal terminology, acronyms, who owns what
Communication craft
quantium-deckPresentations in house style
quantium-documentReports, memos, formal documents, exec briefs
quantium-emailClient and internal correspondence
quantium-visualCharts, diagrams, brand palette
Client engagement
proposal-builderPitches, proposals, RFP responses
sow-and-pricingStatements of work, commercial structuring
discovery-interviewQuestion design and synthesis for client discovery
engagement-planWorkplan, milestones, RAID
client-readoutFindings and final presentations
case-study-writerDelivered work → reusable credentials
pre-send-reviewLast-responsible-moment check before anything leaves the firm
Data and analysis
analysis-planHypothesis, method and test design before touching data
sql-and-queryHouse conventions against internal platforms
notebook-to-narrativeAnalysis → client-readable story
stat-reviewRigour check: significance, confounders, sampling
data-dictionaryDataset documentation
model-cardModel purpose, data, limits, evaluation
Engineering
spec-writerIdea → buildable spec
architecture-decisionADR format
code-reviewHouse review standards
test-strategyCoverage and approach
incident-writeupPostmortems
Risk and governance
ai-risk-assessmentAIMS / ISO 42001-aligned triage for any AI use case
data-handling-checkWhat client data can go where — pre-flight before any tool
contract-reviewAgainst the legal playbook
Internal operations
meeting-to-actionsNotes → decisions, actions, owners
status-update3P-style updates
hiring-packJob specs, interview guides, scorecards

Thirty slots; twelve at launch. The launch set is chosen for highest current duplication and highest cost of being wrong.

Note what is absent: no finance, HR or marketing skills. Those are team level almost by definition — a legitimate difference in how Finance and Marketing work is not a governance failure. Keeping them out is how the org library stays small.

Several of the remaining eighteen should be filled by promotion from community rather than central authoring. That is the better outcome: a promoted skill arrives with an owner who already cares about it.

A.1 Design constraints on specific slots

im-new-here is the largest collision risk in the library. A description like "helps with onboarding and orientation questions" will fire on half of everything. It needs a description written as a hard trigger condition — first-30-days framing, explicit "where do I find" and "how does Quantium do X" phrasings — with negative triggers naming the specific slots it must defer to. Write it last, once the taxonomy is settled, because its job is describing the rest of the library.

Exec briefs are a mode of quantium-document, not a separate slot. They are a genuinely different craft, and they are also close enough that selection between two similar skills will misfire. The failure mode of a wrong pick between two adjacent skills is worse than a slightly fatter skill, so the exec-brief discipline lives inside quantium-document as a mode. This is worth naming explicitly because it is the first slot argument the taxonomy session will have.

A.2 Why there is no fact-check slot

Verification does not belong in the library as a general-purpose skill, for two reasons:

Instead, verification is handled two ways:

  1. Embedded as a mandatory step inside the claim-bearing skills — client-readout, proposal-builder, quantium-document. Every number, date, name and external claim is traced to a source or explicitly flagged as unsourced. No invocation required; it happens at the moment of authorship.
  2. pre-send-review as a genuine slot. "Check this before it goes to the client" is a discrete task with a natural trigger phrase, a real moment in the workflow, and an actively asking user. Verification is one of several things it does, alongside data-handling and confidentiality, commitments accidentally made in prose, figures reconciled against source, and brand and tone.

One binding design constraint on all verification behaviour: a check that asks the model to re-read its own output is close to worthless — self-assessment without new information mostly produces confident re-endorsement. Verification must mandate going outside: web search for external claims, the actual dataset for internal figures, the actual contract for commercial terms. Where no source can be reached, the correct output is a flag, and "unverifiable — flagged" must be a first-class result rather than a failure.