Searchable football data provider and tooling documentation for AI coding agents. Like Context7 for football data.
Who it's for: Developers and analysts who use AI coding tools (Claude Code, Cursor, VS Code Copilot) to work with football data. Works with any tool that supports MCP.
What it does: Gives your AI agent a searchable index of documentation for 23 football data providers and tools — event types, qualifier IDs, coordinate systems, API endpoints, data models, identity surfaces, and cross-provider comparisons for the data providers (StatsBomb, Opta, Wyscout, Impect, SkillCorner, Sportradar, TheSportsDB, FMDB Pro, TransferRoom, and more), plus the open-source libraries people build with (kloppy, mplsoccer, socceraction, soccerdata, floodlight, fast-forward, unravelsports, and more). Your agent looks up the real docs instead of guessing from training data.
Why not just let the AI figure it out? LLMs get football data specifics wrong constantly — Opta qualifier IDs, StatsBomb coordinate ranges, API endpoint URLs, library method signatures. These are mutable facts that change across versions. football-docs gives the agent verified, sourced documentation with provenance tracking so you know where every answer came from.
football-docs is intended to be a community-owned, source-transparent Context7 for football data. The public operating contract is in STRATEGY.md: what belongs here, what must stay out, how we handle public-safe provider facts, and how contributors should prove retrieval quality.
football-docs is the public source for provider identity-surface facts: access shape, ID schemes, matching fields, provider quirks, and provenance rules. Curated provider identity notes belong here when they can be stated as public facts about the provider. They should say whether a fact comes from public docs, public page evidence, licensed feed shape, or a reviewed public-safe observation, and must not include credentials, local paths, internal tooling details, or restricted payloads from any private project.
MCP (Model Context Protocol) is a standard for connecting AI coding tools to external data sources.
claude mcp add football-docs -- npx -y football-docsSettings → MCP → Add server. Use this config:
{
"mcpServers": {
"football-docs": {
"command": "npx",
"args": ["-y", "football-docs"]
}
}
}Add to .vscode/mcp.json:
{
"servers": {
"football-docs": {
"command": "npx",
"args": ["-y", "football-docs"]
}
}
}Add to claude_desktop_config.json:
{
"mcpServers": {
"football-docs": {
"command": "npx",
"args": ["-y", "football-docs"]
}
}
}| Tool | Description |
|---|---|
search_docs |
Full-text search across all provider docs. Filter by provider. Results include provenance (source URL, version). |
resolve_provider_id |
Resolve provider names and aliases to canonical indexed provider keys before searching. |
get_provider_docs |
Retrieve docs for a resolved provider, optionally filtered by topic or category. |
list_providers |
List all indexed providers and their doc coverage. |
compare_providers |
Compare how different providers handle the same concept. |
request_update |
Request a new provider, flag outdated docs, or suggest a better doc source. Queues locally and points to the matching public GitHub issue template. |
resolve_entity |
Resolve players, teams, or coaches to cross-provider IDs via the Reep API. |
Provider filters use the indexed provider keys shown by list_providers, but common aliases are accepted. Examples: fbref, understat, ClubElo, football-data.co.uk, and engsoccerdata search free-sources; Sofascore and ESPN search soccerdata; FMDB searches fmdb-pro; Transfer Room searches transferroom; Hudl Wyscout searches wyscout; Stats Perform / Opta F24 / WhoScored search opta; Metrica, Sportec / DFL, and TRACAB search databallpy; Second Spectrum searches kloppy; Hawk-Eye, SciSports, Signality, Respovision, GradientSports and OptaVision search fast-forward; unravel searches unravelsports; SportRadar API / Soccer Extended search sportradar; The Sports DB / TSDB search thesportsdb; StatsBomb Open Data searches statsbomb.
- "What is Opta qualifier 76?" (big chance)
- "How does StatsBomb represent shot events?"
- "Compare Opta and Wyscout coordinate systems"
- "What player ID fields does Transfermarkt expose?"
- "Does SportMonks have xG data?"
- "What event types does kloppy map to GenericEvent?"
- "How does SPADL represent a tackle?"
| Provider | Chunks | Categories |
|---|---|---|
| fast-forward | 251 | overview, getting-started, data-model, coordinate-system, orientations, layouts, transformations, distributed-compute, api-reference, 12 provider format pages |
| StatsBomb | 235 | event-types, data-model, coordinate-system, api-access, api-endpoints, charting-lineups, xg-model, iq-metrics, player/team stats, player-mapping, identity-surfaces |
| unravelsports | 202 | overview, installation, quickstart, concepts, graph converters, pressing intensity, formation detection, models, utils, american-football |
| Wyscout | 161 | event-types, data-model, coordinate-system, api-access, api-endpoints, charting-analysis-metrics, glossary, identity-surfaces |
| kloppy | 126 | data-model, usage, provider-mapping, tracking-rendering, event-derived-metrics |
| floodlight | 144 | core data objects, io parsers (Tracab, DFL, Kinexon, Opta, SkillCorner, StatsBomb, StatsPerform, Second Spectrum), transforms, metrics, models, visualisation, guides |
| SportMonks | 82 | event-types, data-model, api-access, charting-season-stories, identity-surfaces |
| databallpy | 63 | data-model, overview, usage |
| mplsoccer | 64 | overview, pitch-types, visualizations |
| Impect | 77 | overview, data-model, event-types, coordinate-system, concepts, kpi-definitions, identity-surfaces |
| SkillCorner | 48 | api-access, api-endpoints, data-model, physical-data, coordinate-system, concepts, identity-surfaces |
| Free sources | 62 | overview, fbref, understat, contextual-story-joins, xg-timelines |
| soccerdata | 40 | overview, data-sources, usage |
| TransferRoom | 43 | api-access, api-endpoints, charting-availability, data-model, identity-surfaces |
| Opta | 71 | event-types, qualifiers, coordinate-system, api-access, charting-game-state, charting-lineups, charting-passmaps, charting-set-pieces, charting-shot-placement, identity-surfaces |
| FMDB Pro | 35 | api-access, api-endpoints, data-model, identity-surfaces |
| Sportradar | 29 | api-access, api-endpoints, data-model, charting-and-stories, integration-notes |
| socceraction | 34 | SPADL format, VAEP, Expected Threat |
| BeSoccer | 14 | api-access, api-endpoints |
| TheSportsDB | 18 | api-access, api-endpoints, livescore, identity-surfaces |
| FotMob | 3 | identity-surfaces |
| Soccerdonna | 3 | identity-surfaces |
| Transfermarkt | 3 | identity-surfaces |
1,808 searchable chunks across 23 providers and tools.
Impect documentation is built solely from the public ImpectAPI/open-data repository — a static Bundesliga 2023/24 snapshot, representative of Impect's structure and metric definitions rather than a complete or current mirror. Impect's commercial API is deliberately not documented here. Every enum value, KPI name and field name in
docs/impect/is validated against that repository in CI (pnpm impect:truth,src/__tests__/impect-open-data-validation.test.ts). Data source: Impect; use is subject to the repository's own Terms of Use.
Docs for AI agents are only useful if they are correct, and prose about an API is exactly the kind of thing that drifts or gets invented. Where a machine-readable source of truth exists, this repo checks the docs against it in CI rather than trusting them.
| Providers | Ground truth | Checked by |
|---|---|---|
| kloppy, socceraction, soccerdata, mplsoccer, floodlight, databallpy, skillcorner, fast-forward, unravelsports | The installed package itself — enum members, importable symbols, class constants, Literal parameter vocabularies |
src/__tests__/provider-truth.test.ts |
| Wyscout, SkillCorner, FMDB Pro, Sportradar | The vendor's own publicly published OpenAPI spec — endpoint paths and methods | src/__tests__/provider-truth.test.ts |
| BeSoccer | The vendor's published Postman collection — request vocabulary and parameters | src/__tests__/provider-truth.test.ts |
| Impect | The public open-data repository | src/__tests__/impect-open-data-validation.test.ts |
Truth files live in data/provider-truth/ and are generated, not hand-written:
pnpm provider:truth # rebuild every package truth file (needs python3.11)
pnpm openapi:truth # rebuild every spec-derived truth fileThe specs those derive from are snapshots of publicly published, unauthenticated
vendor documentation. Source URLs, fetch dates and refresh instructions are in
specs/README.md. Wyscout's three API versions merge into one
truth file, because its docs legitimately span v2, v3 and v4.
Each package gets its own pinned venv — co-installing them makes pip silently
downgrade conflicting versions, which would produce truth that disagrees with the
docs. Bump a pin in scripts/gen_all_truth.sh and the matching version in
providers.json together, then re-run and fix whatever the tests flag.
A doc that names an enum member or importable symbol which does not exist in the
real package fails the build. scripts/gen_openapi_truth.py derives the same kind
of facts from a vendor OpenAPI spec, for providers documented that way.
Not every vocabulary is an enum. fast-forward's coordinate systems, orientations
and layouts are lowercase strings on Literal-annotated parameters, so the truth
files also record what each parameter accepts, and a doc writing
coordinates="statsbomb" fails the same way an invented enum member would.
Contributions are welcome from everyone. There are three ways to help:
- Open an issue — request a new provider, flag outdated docs, or suggest a better doc source
- Use the
request_updatetool — AI agents can flag outdated or missing docs directly via the MCP server, which queues requests locally and points to the matching public GitHub issue template - Open a PR — fix errors, add new providers, or improve existing docs
You don't need to be an expert. See CONTRIBUTING.md for the full guide.
Provider doc sources are tracked in providers.json. The crawl pipeline discovers the best doc source (llms.txt > ReadTheDocs > GitHub README) and writes markdown with provenance frontmatter.
npm run discover # probe sources without crawling
npm run crawl # crawl all providers with sources
npm run crawl -- --provider kloppy # crawl one provider
npm run ingest # rebuild search index from docs/
npm run ingest -- --provider kloppy # re-ingest one provider (incremental)Each crawled doc carries provenance metadata (source URL, source type, upstream version, crawl timestamp) that is surfaced in search results, so agents can distinguish between curated content and upstream documentation.
MIT