- Python 99.2%
- Jinja 0.8%
The bot was only ever started by hand, so every reboot left it dead and mon's vibekanka scrape job (METRICS_PORT 9109) read up == 0 from at least 2026-09-02 until it was restarted by hand on 2026-09-23. The playbook installs a system unit running as the checkout's owner, stops a hand-started bot so the two don't fight over the port and Discord token, tightens .env to 0600, and records the mon firewalld rule that was only ever added by hand. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|---|---|---|
| ansible | ||
| data | ||
| docs | ||
| runs | ||
| scripts | ||
| src/vibekanka | ||
| tests | ||
| .env.example | ||
| .gitignore | ||
| CLAUDE.md | ||
| pyproject.toml | ||
| README.md | ||
vibekanka
Turns transcribed Daggerheart session recordings into campaign canon on a self-hosted Kanka-CE instance, with Foundry VTT and GM-authored setting documents as additional ground-truth sources, and a Discord bot in the loop for anything too ambiguous to commit automatically.
Pipeline, per transcript: a session is recorded per-participant and
transcribed upstream (see CLAUDE.md), so speaker labels are ground truth
rather than a diarization guess, then ingested and classified statement-by-statement
into canon facts, ambiguous facts, and out-of-character chatter. Canon facts
are written straight to Kanka-CE. Ambiguous facts get posted as an @mention
in a shared Discord channel for the relevant player or the GM to confirm
before being committed.
Foundry VTT (actors, items, and journal entries) and a GM-provided setting
document can also be synced in as ground truth alongside the transcript
classifier — see CLAUDE.md for the full architecture and design decisions.
Setup
pip install -e ".[dev]"
cp .env.example .env # fill in Kanka-CE, Discord, and LLM credentials
LLM calls go to the Anthropic API by default. To run them against a
self-hosted Ollama instance instead, set
LLM_PROVIDER=ollama plus OLLAMA_BASE_URL/OLLAMA_MODEL — see
.env.example for the context-window caveat, which matters a lot for
full-session transcripts.
Usage
Syncing sources into Kanka-CE:
vibekanka process <transcript> # run one transcript through ingest -> classify -> sync
vibekanka foundry-sync # sync Foundry VTT actors/items/journals into Kanka-CE
vibekanka sync-setting <doc> # parse a GM-provided setting document into Kanka-CE entities
vibekanka sync-clarifications # write clarified (human-answered) facts to Kanka-CE
vibekanka bot # long-running Discord bot for clarification replies
Reviewing and repairing what got synced (all conservative — verify/audit
only report; the merge/rename/reset commands act only on what you name):
vibekanka verify # flag likely duplicate Kanka-CE entities from synced facts, for review
vibekanka audit # full-DB consistency check across all synced sources (add --llm for semantic review)
vibekanka merge-entities # merge two Kanka-CE entities found by verify/audit
vibekanka rename-entity # rename a Kanka-CE entity (e.g. a misheard proper noun)
vibekanka reset-session # undo one session's Kanka-CE writes so it can be reprocessed
vibekanka backfill-subtypes # fill in missing Kanka-CE type sub-categories on existing entities
vibekanka repair-entries # rebuild entity entry text from the local facts table
Proper-noun glossary, used to bias both the classifier and upstream transcription:
vibekanka export-glossary # draft a glossary from live Kanka-CE state, for hand-validation
vibekanka export-whisper-prompt # trim a validated glossary towards WhisperX's 224-token --initial_prompt budget
Choosing a local model (see CLAUDE.md's "LLM providers"), when running classification against a self-hosted Ollama instead of the hosted API:
vibekanka eval-ollama <transcript> # run model(s) N times over a real transcript and report
# per-attempt failure rate, timing, token counts and
# glossary compliance. Makes real LLM calls -- a sweep
# over a local model can take hours.
Development
pytest # run tests
ruff check . # lint
Observability
Logging and Prometheus metrics for everything that leaves the process (LLM
calls with estimated cost, Kanka-CE/Foundry HTTP calls) plus pipeline and
Discord-bot progress counters. vibekanka bot is long-lived and exposes a
normal pull endpoint on METRICS_PORT; one-shot commands instead push to a
Prometheus Pushgateway (PUSHGATEWAY_URL) as they exit, since a short run
usually finishes before a scrape could ever reach it. --metrics-port
overrides both for any command. Both are unset by default.
On lore, vibekanka bot runs as the vibekanka-bot.service systemd unit
(ansible/vibekanka-bot.yml), so it and its METRICS_PORT endpoint come
back after a reboot. Stop the service rather than starting a second bot by
hand; the two would fight over the port and the Discord token.
docs/vibekanka-grafana-dashboard.json is a sample dashboard covering all of
it. ansible/ and docs/*-runbook.md automate the self-hosted Kanka-CE,
Foundry relay, and Pushgateway services this project talks to.
Status
ingestion/loader.py, analysis/classifier.py, kanka/client.py,
storage/db.py, pipeline/orchestrator.py, pipeline/clarifications.py,
pipeline/foundry_sync.py, pipeline/setting_sync.py, pipeline/audit.py,
pipeline/glossary.py, and discord_bot/bot.py are implemented; the
transcript pipeline, Foundry sync, setting-document sync, audit, and glossary
export have all been live-tested end-to-end against real Kanka-CE (and, for
Foundry, a real Foundry VTT instance) data. The Discord bot has likewise been
live-tested against a real Discord server. The Ollama LLM backend is
unit-tested only — there's no Ollama instance on this box to try it against.
See CLAUDE.md for full architecture notes, known gaps, and design decisions.