Star 历史趋势
数据来源: GitHub API · 生成自 Stargazers.cn
README.md

Drama Series Agent

License Python ComfyUI

English · 中文

From a one-line idea to a chained vertical short-drama episode —
develop → write → cast → Ref2VA render — in one series-scoped agent.

https://github.com/user-attachments/assets/52c170e2-bc91-41d2-8b3d-3f2de95c023c

In one series chat you can walk through series settings → episode script → cast portraits → video render, all under the same series_id, then continue into the next episode.


Why this project?

PainWhat we ship
Ideas die in notesskills/ for develop, write, and H3 wiring
Scripts that cannot be shotAuto-built shot tables for ComfyUI (with cast / voice refs)
One-off clips, no shot continuityPrev-shot tail → next first frame — auto chain, no manual editing
Too many separate toolsOne web workbench (+ optional CLI): chat and render progress together

Features

  • What Hermes does (vs ComfyUI)
    ComfyUI + MiniMax H3 turn one shot into video. Hermes covers the path from a sentence to something you can render, resume, and revise:

    1. You say “develop / write EP1” → Hermes writes series settings (plan, world, characters, art style, episode map) and the episode script
    2. Cast portraits are ready → Hermes binds “who appears in which shot” into the shot task table
    3. You say “render EP1” → it builds render files and queues your local ComfyUI
    4. If it stops → resume that episode from the breakpoint on the Jobs page; dislike one shot → re-render only that shot and re-concat
    5. If you change settings or rewrite a script → only the steps that must re-run are marked dirty, not a blind full-episode redo by default

    In short: you should not hand-copy files or manually sync status between chat and ComfyUI.

  • Layered memory (what each layer keeps)

    • Chat log memory/chat.jsonl — what you asked the agent to do
    • Project summary memory/PROJECT.md — series-level facts so later turns still “know which show this is”
    • Series settings dramas/{series}/ — plan, world, characters, art style, episode map (used when writing scripts)
    • Progress state series_runtime.json — how far each episode is (script / cast / render), resume shot, chosen workflow
    • Action trace memory/traces.jsonl — which tools ran (e.g. catch “rewrite script” wrongly starting a render)
      Changing settings or regenerating a script only invalidates the downstream steps that need redo.
  • Where LLMs come from — presets: ModelScope (friendly free tier), Qwen (DashScope), OpenAI, Claude, DeepSeek, Moonshot, local Ollama; any OpenAI-compatible API works

  • Skill packs — develop / write / H3 prompt skills under skills/, editor-agnostic

  • Local render — MiniMax H3 workflows in workflows/selfhost/; you bring ComfyUI + models

  • Cast library — data/cast/{series}/ portraits, optional voice refs

  • Web UI / CLI — chat and jobs in the browser; or drama-series-render for a headless episode

Hermes workbench main UI

Workbench: chat on the left, per-episode progress and render jobs on the right


Quick start

Requirements

  • Python 3.10+
  • ffmpeg (see Install ffmpeg below)
  • Optional: Node 18+ for the React workbench
  • Optional: OpenAI-compatible LLM for the agent chat (e.g. ModelScope)

ComfyUI (required for video render — install yourself)

This repo does not bundle ComfyUI or MiniMax H3 weights. You must:

  1. Install ComfyUI yourself and keep it running (default http://127.0.0.1:8188).
  2. Install the MiniMax H3 Ref2VA custom nodes your graphs need (e.g. MiniMaxH3ReferenceToVideo and related loaders / optional SageAttention · BlockCache nodes used by the *_fast / *_turbo tiers).
  3. Download every model file referenced by the workflow JSON into the matching ComfyUI folders (models/diffusion_models, models/text_encoders, models/vae, models/loras, … — follow ComfyUI’s usual layout). Typical names from workflows/selfhost/:
File (examples)Used by
minimax_h3_ref2va_pruned_int8_convrot.safetensorsUNET / Ref2VA (fast · lora · turbo)
minimax_h3_fl2va_pruned_int8_convrot.safetensorsQuality tier video_minimax_h3_r2v.json
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensorsText encoder / CLIP
minimax_h3_video_vae_fp16.safetensorsVideo VAE
minimax_h3_audio_vae_fp32.safetensorsAudio VAE
minimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16.safetensorsLightning LoRA (*_lora / *_turbo)
  1. Open the chosen workflow once in the ComfyUI UI and confirm no red/missing nodes or models, then point this agent at the same server (COMFYUI_URL / system config).

Without a ready ComfyUI + models, you can still develop settings, write scripts, and manage cast in the workbench — video jobs will fail until the above is done.

Install ffmpeg

ffmpeg is required after each shot to extract the tail frame (chain continuity) and to concat shots into master.mp4.
Resolution order in code: FFMPEG_PATH → PATH → bundled imageio-ffmpeg (pulled in by pip install).

Recommended: install a system binary (clearer for debugging):

Windows
# Option A — winget
winget install --id Gyan.FFmpeg -e

# Option B — scoop
scoop install ffmpeg

# Option C — chocolatey
choco install ffmpeg

Or download a build from gyan.dev ffmpeg builds / BtbN, unzip, and either:

  1. Add ...\ffmpeg\bin to User PATH, or
  2. Set in .env:
FFMPEG_PATH=C:\ffmpeg\bin\ffmpeg.exe

Verify in a new terminal:

ffmpeg -version
macOS
brew install ffmpeg
ffmpeg -version
Linux (Debian/Ubuntu)
sudo apt update
sudo apt install -y ffmpeg
ffmpeg -version

Fallback (no system ffmpeg): pip install -e . already depends on imageio-ffmpeg, which downloads a static binary into the venv. That is enough for Quick Start; set FFMPEG_PATH only if you want to force a specific build.

Install the project

git clone https://github.com/binwu1/drama-series-agent.git
cd drama-series-agent

python -m venv .venv
# Windows:
.venv\Scripts\activate
# macOS/Linux:
source .venv/bin/activate

python -m pip install -U pip
pip install -e ".[render,dev]"
copy .env.example .env          # Windows
# cp .env.example .env          # macOS/Linux
# edit .env → COMFYUI_URL / LLM_* / optional FFMPEG_PATH

Verify install

python scripts/doctor.py

You should see [OK] for ffmpeg, skills, comfykit, API router, and CLI entry points.
If ffmpeg fails: install a system binary (above) or confirm imageio-ffmpeg is in the active venv.

Launch API (+ web workbench)

# Windows — API + Vite UI
.\start.ps1 -WithWeb

# macOS / Linux
./start.sh --with-web
SurfaceURLNotes
Web workbenchhttp://localhost:5173/Default Vite port (web/). Chat, series settings, scripts, cast, jobs, system config.
API (OpenAPI)http://127.0.0.1:8000/docsBackend the UI calls
API base used by UIhttp://127.0.0.1:8000/api/seriesOverride with web/.env: VITE_HERMES_API_BASE=...

Open http://localhost:5173/ after both processes are up. If the page loads but chat/settings fail, the API is not running on port 8000 (start it first, or point VITE_HERMES_API_BASE at your API).

Without the helper scripts:

# terminal 1 — API
python scripts/run_api.py --host 127.0.0.1 --port 8000

# terminal 2 — workbench (http://localhost:5173/)
cd web
npm install
npm run dev

Render one episode (CLI)

  1. Cast: data/cast/{series}/{character}.png (+ optional voices/{character}.wav)
  2. Run files: templates/{series}/run/EP001.episode-run.jsonl + EP001.episode-meta.json
  3. ComfyUI at COMFYUI_URL (default http://127.0.0.1:8188)
drama-series-render \
  --project templates/<series> \
  --episode EP001 \
  --cast-series <series> \
  --context-ir off \
  --comfyui-url http://127.0.0.1:8188

Output: templates/{series}/output/{EP}/master.mp4 (shots concat with a short head trim).


Workflows

API graphs live under workflows/selfhost/. Import / align them in your ComfyUI, and download any model the JSON names (see ComfyUI requirements).

WorkflowRole
selfhost/video_minimax_h3_r2v_fast.jsonDefault speed tier (20 step + accel nodes)
selfhost/video_minimax_h3_r2v_lora.jsonLightning LoRA 4-step path
selfhost/video_minimax_h3_r2v_turbo.jsonTurbo / experimental speed
selfhost/video_minimax_h3_r2v.jsonQuality / fl2va tier

Pick the video workflow in the system config panel or config.json → comfyui.video.default_workflow.


Repository layout

skills/                         # canonical agent skills (open-source layout)
src/drama_series_agent/
  agent/  drama/  intake/  api/ # orchestration
  adapters/r2v/                 # H3 episode runner
workflows/selfhost/             # ComfyUI API graphs
data/cast/{series}/             # character anchors + voices/
projects/{series}/              # series_runtime, memory, jobs
templates/{series}/             # run/ + output/
web/                            # React workbench
docs/                           # architecture, contracts, assets

Cursor users who want project-skill auto-discovery can run scripts/link_cursor_skills.ps1 / .sh (links are gitignored). Edit skills only under skills/.


Documentation


Contributing

Issues and PRs are welcome. Keep changes focused; prefer small PRs with a clear why.

  1. Fork & branch from main
  2. pip install -e ".[render,dev]"
  3. Open a PR with a short summary and test notes

License

Apache License 2.0

Literary skill content under skills/0xsline-short-drama/ retains its upstream license notice where present.

关于 About

一句话出片的短剧系列 Agent(Series-scoped agent: idea → short-drama script → cast → MiniMax H3 Ref2VA render)

语言 Languages

Python79.3%
TypeScript16.6%
CSS2.8%
PowerShell0.8%
Shell0.6%
HTML0.1%

提交活跃度 Commit Activity

代码提交热力图
过去 52 周的开发活跃度
6
Total Commits
峰值: 4次/周
Less
More

核心贡献者 Contributors