Wiki to Video

How prompts and videos are built

Cursor is the editor this app was written in. It does not generate scripts at runtime. Keep this page in sync with README.md when the pipeline changes.

What AI is used

The only writing models are optional OpenAI Chat Completions (default gpt-4o-mini; also gpt-4o, gpt-4.1-mini, gpt-4.1) local Ollama, or Claude from Admin, and they only rewrite the topic brief (hooks, topics, facts). They must not invent facts. The spoken script, scene board, and InVideo prompt are Python templates in prompts.py — no second LLM call. Agnes (or local LTX) is image-to-video, not a writer. Voice is Edge TTS (Male / Female / Warrior / Ancient, plus named Andrew, Jenny, Aria, Guy, Christopher, Emma, Sonia, Ryan).

PythonScripts and InVideo prompt
OpenAI / Ollama / ClaudeOptional brief rewrite
Agnes I2V / T2VMotion from still or text

Runtime path

  1. Source (no writing AI)

    Wikipedia fetch, topic search, or a pasted script. Python only. If Britannica or another allowlisted host returns 403, Build uses the English Wikipedia page from the URL slug.

  2. Brief (optional OpenAI, Gemini, Groq, Claude, or Ollama)

    Python extracts sentences. The script model for this Build can rewrite hooks and topics toward what/why/wow (not the same establishment year restated), and for listicle topics it should name several list items — not only the survivor. New prompt is pre-filled from Admin; changing it does not save Admin. Default is OpenAI gpt-4o-mini when OPENAI_API_KEY is set (also gpt-4o, gpt-4.1-mini, gpt-4.1). Gemini and Groq are free-tier cloud (not local GPU). Gemini uses generateContent when GEMINI_API_KEY is set (GOOGLE_API_KEY also works): default gemini-3.6-flash (free-tier), plus gemini-3.5-flash-lite and gemini-2.5-pro. Get a Gemini key at https://aistudio.google.com/apikey. Groq is console.groq.com (not xAI Grok) when GROQ_API_KEY is set: default openai/gpt-oss-20b (free-tier), plus openai/gpt-oss-120b and qwen/qwen3.6-27b. Get a Groq key at https://console.groq.com → sign in → API Keys → Create API Key. Restart uvicorn after adding a key. OpenAI and Claude are paid. Ollama is free-local but can freeze this machine; it uses a model already pulled (OLLAMA_HOST, default http://127.0.0.1:11434). Claude uses the Anthropic Messages API when CLAUDE_API_KEY is set (claude-sonnet-5 for everyday scripts; claude-haiku-4-5-20251001 is cheaper; claude-opus-5 is expensive). A selected provider never silently calls another. If it is not ready, Build keeps the extractive brief and shows a warning. New Ollama models need ollama pull <name> — this app will not pull for you. Wikidata/Grokipedia stay on the brief, not in voice or burned-in captions.

  3. Platform scripts (Python templates)

    prompts.py fills TikTok / Reels / Shorts / YouTube: voiceover, scenes, InVideo text, Veo shots. Short-form beats prefer one clear hook plus varied what/why/wow facts and skip repeating the same establishment year across captions. Listicle topics (Seven Wonders, named catalogs) pack item names into 1–2 beats, avoid repeating the survivor claim, align on-screen chips to the beat when clear, and allow one strong year punch once. Tone Animated is comics / anime 2D motion (Agnes, then a still if a clip misses — Agnes never loads LTX); Explainer is motion-graphics / here's-why; Mythbust is claim then record. Other tones stay faceless documentary. No second LLM call. YouTube long-form builds ~8 minutes of spoken script from section bodies (~150 wpm) with the same date-lead dedupe. Build also assigns still URLs to each scene.

  4. Save + queue

    One SQLite file in the project folder for all prompts and jobs. Scene-board stills stay in TEMP_DIR until the first job, then the prompt folder moves to output with a scratch folder per platform (preview /media/ links follow that move). After Build, Replace a still from a local jpg/png/webp (copied into that folder). Queue defaults to CPU preview (Pexels stock video when a beat matches, else Ken Burns — no Agnes, no Local GPU). When that pack looks right, Upgrade pack to Agnes from Jobs reuses the same script, voice, and stills. Two jobs run at once. Only one Agnes job at a time; queued Agnes jobs wait. GPU and stills can still run in parallel with each other. A second Local GPU job waits until the 3060 is free.

  5. Render (not prompt AI)

    Edge TTS speaks the script. Queue mixes a quiet Openverse instrumental (CC0 / public domain / CC BY, no API key) under the voice. Stills prefer on-article Wikipedia/Wikimedia images, including fair-use title cards (skip icons, tiny frames, cosplay, graffiti, and merch). Documentary tones rank ruin / ash / crater / excavation stills over scenic harbor shots when the topic or voiceover is about death, eruption, or archaeology, and assign each beat by matching the scene line to the filename — not only by pool index. Then AniList posters (no key) and optional TMDB if TMDB_API_KEY is set — catalog key art first for Animated / fiction at Build; Educational and Mythbust stay Wikipedia + documentary stock (no catalog art) — then Wikimedia Commons search if the page is thin. Openverse fill uses beat-aware queries and the same ranking; a hit must mention a distinctive title word or it is dropped (so “TV series” does not pull Netflix posters). Optional Pexels stills only for non-fiction pages if PEXELS_API_KEY is set. Animated and fiction titles skip Pexels. Rebuild re-fetches unlocked remote stills; Lock and local files stay. You can Skip still to cycle a beat to another candidate, Lock a pick, Replace from a local jpg/png/webp (copied into the pack folder), or Add still from disk for the pack. Queue / CPU preview use those chosen stills in scene order. For documentary tones, CPU preview also searches Pexels video per beat (same subject ranking: ruin/ash over harbor), trims the clip to the voiceover slot, and Ken Burns the still only when no clip matches or the clip is shorter than the slot — no looping. Agnes stays still → image-to-video (stock video is preview-only). Animated / fiction / explainer skip stock video. Agnes swaps Wikimedia scene stills per clip to AniList/TMDB/Pexels (fetchable hosts) so create does not hang on rehost-only sets. Locked local stills are kept. Wikimedia is rehosted only when no fetchable substitute exists. Mixing fetched and local is fine. Explainer prefers diagrams / infographics. Mythbust keeps documentary stills but labels beats THE CLAIM / THE RECORD. YouTube long-form uses ~12 stills. Agnes or local GPU animate ~5s clips (play once, then Ken Burns). Short-form uses ~10 stills by default (up to 12 on longer voiceovers) so slots stay near 5s and Ken Burns tails stay short. Queue offers Agnes image→video (still + ti2vid) or Agnes text→video (prompt only, no still sent; filename -agnes-t2v). T2V briefs lock onto a primary subject from the voiceover (listicle beats film one named item, not a collage). Motion prompts pick a per-beat camera (push-in, orbit, parallax, or energy) from each scene’s visual line. Fiction / Animated packs prefer AniList/TMDB key art on the scene board at Build, and skip weak commons (cosplay, graffiti, merch). Listicle shorts speak one named item per beat (not a full catalog dump) so captions and motion stay on the same subject. Scene stills are matched to that name (Commons fill when the page gallery is thin); Agnes keeps those stills and rehosts Wikimedia instead of swapping in unrelated stock. Agnes defaults to agnes-video-v2.0 (Admin can keep that or show Video 2.5 as not I2V-ready yet). Not the 2.5 text models. Free AGNES_API_KEY only. PAYED_AGNES_KEY is ignored until free is stable. Create waits 300s for a task id (one float for all HTTP phases). A timeout on one clip uses a still and the next clip still tries free Agnes. The job fails at the end only if no clip got motion. HTTP 502 / 503 / 520 (including video_queue_full) wait ~20s and retry the same free key up to twice more (3 tries total), then that clip is a still. Agnes jobs prefer Pexels/Openverse CDNs and rehost Wikimedia before create. The header Agnes chip is green when their API is reachable, red when it is full or timing out. While nothing is queued or running, the app re-checks about once an hour with GET /v1/models (not a fake video create). ffmpeg stitches the MP4. Portrait captions sit above TikTok/Reels chrome (~5-word phrases) in a local OFL font (Auto from topic — history for volcanoes / ruins — or Queue override). YouTube long-form leaves the picture clean and writes an SRT next to the MP4. Duplicate spoken lines and mid-sentence fragments are dropped. Local LTX keeps T5 on CPU, the 2B transformer on the 3060, and the VAE in float32 (3D conv has no CUDA bfloat16 kernel). Queue GPU quality: Draft, Balanced (default, 41 frames at 512×768), or High. Ken Burns fills the rest of the slot; ffmpeg uses CRF 18 on Balanced. VRAM is capped so the display GPU does not freeze.

Diagram

New prompt
  ├─ Topic search ──► Wikipedia / Grokipedia / allowlisted sites ──► Fetch best source
  ├─ Wikipedia URL ───────────────────────────────────────────────► Fetch Wikipedia
  └─ Pasted external prompt ──► Parse text in Python ─────────────► Topic brief
Fetch Wikipedia ──► Extractive brief (Python) ──► Enrich? ──► Topic brief
Topic brief ──► Script provider ready? (this Build, else Admin)
                 ├─ OpenAI key ──► this Build's OpenAI model rewrites brief only
                 ├─ Gemini key ──► this Build's Gemini model rewrites brief only
                 ├─ Groq key ────► this Build's Groq model rewrites brief only
                 ├─ Claude key ──► this Build's Claude model rewrites brief only
                 ├─ Ollama up ───► this Build's Ollama model rewrites brief only
                 └─ no ──────────► prompts.py templates
prompts.py ──► Save pack as a row in project SQLite (stills in TEMP_DIR)
Save ──► Queue job (engine / tone / voice / captions) ──► move prompt folder to OUTPUT_DIR ──► Edge TTS + ranked stills + Openverse music
                    ├─ Agnes / GPU ──► ~5s motion, then Ken Burns ──► ffmpeg
                    └─ CPU preview ──► Pexels stock video (else Ken Burns) ──► ffmpeg ──► prompt-folder/filename.mp4
Jobs ──► Upgrade pack to Agnes (same script / voice / stills; one Agnes job at a time)
YouTube long-form also writes filename.srt (no burned-in captions).
Two jobs run at once. A second Local GPU job waits for the 3060.

Who writes what

Piece Who produces it Model / tool
Wikipedia / search / paste This repo (Python fetch or your paste) None
Extractive brief summarize.py sentence scoring None
Rewritten hooks / topics Only if this Build's Script provider is ready (OpenAI, Gemini, Groq, or Claude key, or Ollama up with a pulled model). New prompt is pre-filled from Admin. This Build's OpenAI / Gemini / Groq / Claude / Ollama model (stamped on the pack and jobs)
Spoken script, scenes, InVideo prompt prompts.py templates. Short-form skips repeating the same establishment year, packs listicle item names into 1–2 beats, dedupes survivor claims, aligns chips when clear, and allows one strong year punch. Prefers varied what/why/wow beats. YouTube long-form fills ~1100–1300 spoken words (~8 min at 150 wpm) from section bodies with date-lead dedupe; timestamps follow word count. Short articles stay short. No silence padding. None
Voiceover audio Edge TTS Edge TTS (roles + named: Andrew, Jenny, Aria, Guy, Christopher, Emma, Sonia, Ryan)
Background music Openverse audio search (instrumental, CC0 / pdm / CC BY). Duck under speech. Optional: skip if the download fails. None
Stills On-article Wikimedia (incl. fair-use; skip cosplay/graffiti/merch); beat-matched to voiceover; AniList/TMDB for Animated/fiction only (Educational stays documentary); Openverse fill must match the title; Pexels only for non-fiction. Local replace for episode frames. None
Motion clips After you queue a job. CPU preview uses Pexels stock video when available, else Ken Burns; upgrade the same pack to Agnes from Jobs. Agnes agnes-video-v2.0 by default (Admin; 2.5 listed but not I2V yet) or local LTX
Final MP4 ffmpeg mux + OFL captions (SRT on YouTube long-form) None

Start the server

Run this in a GNOME terminal, not inside Cursor (closing Cursor kills that process). Skip --reload so a long Agnes or Local GPU job is not killed by a code save.

cd ~/wiki_to_video
source .venv/bin/activate
uvicorn app:app --host 0.0.0.0 --port 8765

Open http://127.0.0.1:8765

Operator flow

  1. New prompt — search a topic (Build uses the first source unless you pick another), paste a URL, or paste an external script. Script / LLM is pre-filled from Admin; change it for this Build only.
  2. Build — always rewrites the brief with the Script / LLM selected on New prompt (this Build only), even if a pack with that title already exists. Wikipedia can be fetched again; local stills are kept when the scene list lines up. Stills stay in temp until you queue a job. Same title + tone + platforms overwrites the pack, not old Jobs rows.
  3. Edit script — change voiceover, on-screen text, visuals, or caption. Each beat shows the still Queue will use; Skip still cycles to another candidate, or lock before Queue. Agnes waits until Queue.
  4. Queue — moves the prompt folder to output, then the Admin default engine (shipped as CPU preview), Agnes, or local GPU, plus a quiet Openverse instrumental under the voice. You can still override the engine per job. CPU preview uses Pexels stock video clips when a beat matches (else Ken Burns). Tone and Voice are the same dropdowns as New prompt. Queue tone changes look, music, still preference, and filename; script shape stays from Build. Captions picks an OFL font for TikTok / Reels / Shorts (Auto from topic); YouTube long-form writes an SRT instead of burning text. Two jobs run at once (separate SQLite connections so parallel queue writes cannot crash a job). Agnes uses the free AGNES_API_KEY unless Admin fallback is free→paid and PAYED_AGNES_KEY is set. The model defaults to agnes-video-v2.0 (change on Admin; Video 2.5 is listed but disabled). Create waits 300s for a task id. 503 retries stay on the current key. A miss stays a still. Only one Agnes job at a time (queued Agnes jobs wait). GPU and stills can still run in parallel with each other. A second Local GPU job waits for the 3060. Local GPU has Draft / Balanced / High. Each platform gets a scratch folder; the MP4 lands at the prompt root.
  5. Jobs — started, ended, duration, play MP4, copy caption/hashtags, logs, retry failed, output filename (not the host path). CPU preview rows show a preview badge. Expanded rows (and a small badge on the row) show which script model wrote the brief, stamped at Build — e.g. Groq · openai/gpt-oss-20b, or extractive (no LLM) when the rewrite did not run. Upgrade pack to Agnes reuses the same script, assets, and brief LLM stamp (separate Jobs rows, not a retry). The Agnes column is Free. Retry queues a new job; the table shows one row per chain with attempt history.
  6. Admin — settings page. Pick the Script provider (OpenAI, Ollama, Claude, Gemini, or Groq) and model (OpenAI default gpt-4o-mini; also gpt-4o, gpt-4.1-mini, gpt-4.1; Ollama list comes from this machine; Claude: claude-haiku-4-5-20251001, claude-sonnet-5, claude-opus-5; Gemini: gemini-3.6-flash default free-tier, gemini-3.5-flash-lite, gemini-2.5-pro; Groq: openai/gpt-oss-20b default free-tier, openai/gpt-oss-120b, qwen/qwen3.6-27b). Gemini and Groq are free-tier cloud; Ollama is free-local but can freeze the machine; OpenAI and Claude are paid. Also the Agnes video model (default agnes-video-v2.0), free-only vs free→paid fallback, default Queue engine (CPU preview / Agnes / Local GPU), and Light / Dark theme. Shows OpenAI / Claude / Gemini / Groq key present/missing, Ollama reachable/not running, paid-key present/missing, and last Agnes GET /v1/models probe (video models + queue note). Saved in SQLite. Theme applies to every page immediately. Keys are not shown.

Admin and Agnes Video 2.5

Admin stores these settings in SQLite: Script provider (openai default, ollama, claude, gemini, or groq) and model (OpenAI: gpt-4o-mini default, plus gpt-4o, gpt-4.1-mini, gpt-4.1; Ollama: tags from GET /api/tags on OLLAMA_HOST; Claude: claude-haiku-4-5-20251001, claude-sonnet-5, claude-opus-5; Gemini: gemini-3.6-flash default, plus gemini-3.5-flash-lite, gemini-2.5-pro; Groq: openai/gpt-oss-20b default, plus openai/gpt-oss-120b, qwen/qwen3.6-27b), Agnes model, fallback (free or free_then_paid), default Queue engine (stills / agnes / gpu), channel branding (name, YouTube / Instagram / TikTok @ handles, logo path for watermark), and page theme (light is the current default; dark is a full page theme). Renders burn the logo as a watermark on every clip and show the channel name plus the platform @ on the last few seconds (follow-for-more beat). Default logo: static/brand/history-buffer-icon.png. New Queue jobs use the saved engine when you do not pick one. New Agnes jobs stamp the selected model at queue time. Fallback default is free (current proven mode). free_then_paid adds the paid slot only when PAYED_AGNES_KEY is present; if the key is missing, jobs keep using free and Admin shows a warning. The key itself is never shown. Agnes status is a GET /v1/models probe (refresh on the page) — video models and last queue/health note, no billed create.

Publishing: connect YouTube on Admin, then Publish to YouTube on done long-form or Shorts jobs (unlisted). TikTok and Instagram Reels are manual uploads from the NAS exports; copy caption from Jobs. Use the Instagram MP4 for Facebook Reels too (same 9:16 file — upload to the Page or cross-post from Meta Business Suite). Backlog: mark-published links for TikTok / IG / Facebook, then API automation. YouTube PT: Portuguese captions (captions.insert) and dubbed audio (PT-BR voiceover file for Studio Languages — API does not upload dubs yet). Channel bundles: multiple named channel profiles (logo, handles, YouTube OAuth per bundle) instead of one global History Buffer set — pick on Queue.

Video 2.5 evaluation (17 Aug 2026): official docs for agnes-video-2.5 are a pre-release (agnes-video-v25). The create body is a different Videos API (mode is text / keyframe / reference, first_frame instead of image + ti2vid; width / height / num_frames return 400). Poll is GET /v1/videos/{id}, not /agnesapi. A live GET /v1/models with the free key listed only agnes-video-v2.0 among video models (plus text and image models we do not offer). So 2.5 is shown on Admin as not I2V-ready yet. Default stays V2.0.

Full operator notes: README on GitHub and README.md in the repo.