Turn manhwa & webtoon chapters into narrated recap videos โ fully automated.
AniFlow downloads manhwa/webtoon chapters, detects and crops panels using YOLO, extracts speech-bubble text, generates AI narration, converts it to speech, and renders video clips โ all through a browser-based UI.
- ๐ Web Scraper โ download chapters from manhuaus, manhwatop, MangaDex, and more
- ๐ YOLO Panel Detection โ custom-trained models detect panels and speech bubbles
- ๐งน Text Removal โ automatically clean speech bubbles from art panels
- ๐ค AI Narration โ local (Ollama) or cloud (Gemini / Claude) narration
- ๐๏ธ Multi-TTS โ Edge TTS (free), Kokoro (local), or ElevenLabs (cloud)
- ๐ฌ Video Rendering โ Ken Burns animation, blur backgrounds, BGM, watermarks
- ๐ฅ Character Casting โ import from MyAnimeList or scan faces from chapters
- ๐ฆ Batch Mode โ process dozens of chapters unattended
| Dialogue Recap | Narrated Recap | |
|---|---|---|
| Panels | Original panels with speech bubbles | Clean art panels (bubbles removed) |
| Audio | Extracted dialogue read aloud | AI-written YouTube-style recap narration |
| Best for | Letting the original dialogue shine | YouTube story recaps |
๐ See it in action! Check the
example_output/folder for sample clips from both modes.
| Tutorial | Link |
|---|---|
| ๐ฌ Dialogue Recap โ Single Chapter | Watch on YouTube |
| ๐ฌ Dialogue Recap โ Batch Mode (50+ Chapters) | Watch on YouTube |
| ๐๏ธ Narrated Recap โ Automatic Pipeline | Watch on YouTube |
| ๐๏ธ Narrated Recap โ Batch Mode (Full Series) | Watch on YouTube |
| Dependency | Purpose | Install |
|---|---|---|
| Python 3.10 โ 3.12 | Runtime | python.org |
| FFmpeg | Video/audio processing | brew install ffmpeg (macOS) ยท sudo apt install ffmpeg (Linux) ยท ffmpeg.org (Windows) |
| 9router / Ollama | Local AI narration & text extraction | 9router / ollama.com |
| ~10 GB disk | Python env + model weights | โ |
GPU recommended but not required. Apple Silicon Macs and NVIDIA GPUs will significantly speed up panel detection and Kokoro TTS.
git clone https://github.com/aashish254/Aniflow.git
cd AniFlow
chmod +x setup.sh
./setup.shThe setup script will check dependencies, create a virtual environment, install packages, pull Ollama models, and start the app.
# 1. Clone the repository
git clone https://github.com/aashish254/Aniflow.git
cd AniFlow
# 2. Create a virtual environment (Python 3.12 recommended)
py -3.12 -m venv venv # Windows
# python3 -m venv venv # macOS / Linux
# 3. Activate virtual environment
# Windows (PowerShell):
.\venv\Scripts\Activate.ps1
# macOS / Linux:
# source venv/bin/activate
# 4. Install Python dependencies (~10-20 min, torch is large)
pip install --upgrade pip
pip install -r requirements.txt
# 5. Configure 9router (Alternative to Ollama)
# If you use 9router, set these environment variables before running:
# Windows (PowerShell):
# $env:USE_NINE_ROUTER="1"
# $env:NINE_ROUTER_HOST="http://127.0.0.1:20128"
# $env:VISION_MODEL="Free-Tier"
# $env:WRITER_MODEL="Free-Tier"
# $env:TEXT_WRITER_MODEL="Free-Tier"
# 6. Pull Ollama models (Skip if using 9router)
# ollama pull qwen2.5vl:3b
# ollama pull qwen2.5:14b
# 7. Run the app
python app.py
# 8. Open UI in browser:
# http://localhost:8080| Dependensi | Fungsi | Instalasi |
|---|---|---|
| Python 3.10 โ 3.12 | Runtime | python.org |
| FFmpeg | Pemrosesan video/audio | brew install ffmpeg (macOS) ยท sudo apt install ffmpeg (Linux) ยท ffmpeg.org (Windows) |
| 9router / Ollama | AI lokal untuk narasi & ekstraksi teks | 9router / ollama.com |
| ~10 GB disk | Python env + model weights | โ |
GPU disarankan tapi tidak wajib. Apple Silicon Mac dan GPU NVIDIA akan sangat mempercepat deteksi panel dan Kokoro TTS.
Pastikan 9router sudah berjalan di latar belakang (port default: http://127.0.0.1:20128).
# 1. Masuk folder proyek
cd D:\Aniflow
# 2. Aktifkan virtual environment
.\venv\Scripts\Activate.ps1
# 3. Set environment variables untuk 9router
$env:USE_NINE_ROUTER="1"
$env:NINE_ROUTER_HOST="http://127.0.0.1:20128"
$env:VISION_MODEL="Free-Tier"
$env:WRITER_MODEL="Free-Tier"
$env:TEXT_WRITER_MODEL="Free-Tier"
# (Opsional) Jika 9router butuh API Key
# $env:NINE_ROUTER_API_KEY="API_KEY_ANDA"
# 4. Jalankan server AniFlow
python app.pyAkses antarmuka web di:
http://localhost:8080
setup.bat
Or follow the manual steps above using `venv\Scripts\activate` instead.
---
## ๐ฌ How to Use
### Narrated Recap (Recommended)
1. **Source** โ Paste a manhwa URL โ **Analyze** โ pick chapters โ **Download**
(or point to a local folder of images)
2. **Cast** *(optional)* โ Import characters from MyAnimeList, or scan chapter faces
and assign names. The narrator uses real names ("Jin-Wooโฆ") instead of "the dark-haired guy."
3. **Stitch** โ Stitch โ Detect (bubble panels) โ Crop โ Stitch โ Detect (art panels) โ Crop.
Boxes are editable โ drag, resize, or draw new ones.
4. **Review** โ Check cropped panels; delete junk or re-crop.
5. **Extract** โ AI reads speech-bubble text.
6. **Narrate** โ Choose your AI backend:
- ๐ฆ **Ollama (local)** โ free & offline. `qwen2.5vl:7b` (vision) or `qwen2.5:14b` (text-only).
- โจ **Gemini API** โ paste multiple API keys (auto-rotation on quota). Free keys at [aistudio.google.com](https://aistudio.google.com).
- ๐ค **Claude API** โ single key.
7. **Review Text** โ Edit any narration line or re-narrate individual panels.
8. **Audio** โ TTS with Edge (free), Kokoro (local), or ElevenLabs.
9. **Video** โ Renders clips with Ken Burns animation, blur backgrounds, BGM.
10. **Review Clips / Sort Clips** โ Preview, download, or batch-merge.
### Dialogue Recap
Same flow minus the cast/narrate steps โ extracted dialogue becomes the audio,
panels keep their original speech bubbles.
### Batch Mode
Download chapters 1โ68, enable **Batch Mode**, hit **Run Full Batch** โ every step
runs across all chapters automatically. Progress is streamed live.
---
## ๐ Project Structure
AniFlow/ โโโ app.py # Flask web app + API routes โโโ config.py # All configuration & path settings โโโ requirements.txt # Python dependencies โโโ setup.sh / setup.bat # One-command setup scripts โโโ start.sh # Quick-start script (macOS/Linux) โ โโโ pipeline/ # Core processing modules โ โโโ web_scraper.py # Chapter downloading (multi-site) โ โโโ stitcher.py # Image stitching โ โโโ panel_extractor_yolo.py # YOLO panel detection โ โโโ cropper.py # Panel cropping โ โโโ text_extractor.py # OCR / text extraction โ โโโ text_remover.py # Speech bubble cleaning โ โโโ narrator.py # Local AI narration (Ollama) โ โโโ gemini_narrator.py # Cloud narration (Gemini API) โ โโโ scene_narrator.py # Scene-level narration โ โโโ tts_engine.py # Edge TTS engine โ โโโ kokoro_tts.py # Kokoro local TTS โ โโโ elevenlabs_tts.py # ElevenLabs cloud TTS โ โโโ compositor.py # Video rendering โ โโโ orchestrator.py # Pipeline orchestration โ โโโ cast_system.py # Character management โ โโโ character_builder.py # Face detection & grouping โ โโโ mal_character_importer.py # MAL character import โ โโโ script_rewriter.py # Narration style rewriter โ โโโ capcut_director.py # CapCut export โ โโโ panel_filter.py # Panel quality filtering โ โโโ models/ # YOLO model weights (committed) โ โโโ manhwa_panel_cropped.pt # Art panel detector โ โโโ panel_with_bubble_text.pt # Bubble panel detector โ โโโ templates/ # HTML templates โโโ static/ # CSS, JS, logos โโโ assets/ # BGM, watermarks, intros, outros โโโ tests/ # Test & diagnostic scripts
### Generated Folders (gitignored)
| Folder | Contents |
|---|---|
| `manhwa_download/` | Downloaded chapter images + per-chapter state |
| `manhwa_stitch/` | Stitched full-page strips |
| `manhwa_detect/` | Detection metadata (bounding boxes) |
| `manhwa_crop/` | Cropped individual panels |
| `manhwa_text/` | Extracted dialogue JSON |
| `manhwa_audio/` | Generated TTS audio files |
| `manhwa_clips/` | Final rendered video clips |
| `output/` | Legacy output directory |
| `~/RECAP/Cast Data/` | Your character casts (faces + profiles) |
All generated data is **regenerable** โ delete any folder and re-run the pipeline.
---
## โ๏ธ Configuration
### API Keys (Optional)
Copy `.env.example` to `.env` and fill in keys only for services you want to use:
```bash
cp .env.example .env
| Key | Service | Required? |
|---|---|---|
GEMINI_API_KEY |
Google Gemini narration | Only if using Gemini |
ANTHROPIC_API_KEY |
Claude narration | Only if using Claude |
ELEVENLABS_API_KEY |
ElevenLabs TTS | Only if using ElevenLabs |
MAL_CLIENT_ID |
MyAnimeList import | Only if importing characters from MAL |
Ollama + Edge TTS work completely offline โ no API keys needed for the core pipeline.
| Model | Role | Size | Required? |
|---|---|---|---|
qwen2.5vl:3b |
Vision parser (reads panels) | ~2 GB | โ Required |
qwen2.5:14b |
Text-only narrator | ~9 GB | ๐ Recommended |
qwen2.5vl:7b |
Vision narrator (sees panels + cast) | ~5 GB | Optional |
| Problem | Solution |
|---|---|
| "Address already in use" | Another instance is running: lsof -ti :8080 | xargs kill -9 |
| Ollama dot is red | Start Ollama and pull models: ollama pull qwen2.5vl:3b |
| Downloads hang | Protected site โ the app auto-switches to headless browser (~40s). MangaDex URLs are most reliable. |
| Bubbles not removed | EasyOCR inpainter handles clear bubbles; complex ones can be fixed with the crop tool in Review. |
ModuleNotFoundError |
Make sure your venv is activated: source venv/bin/activate |
| Torch install fails | Try: pip install torch torchvision --index-url https://download.pytorch.org/whl/cpu |
AniFlow is built on the shoulders of many incredible open-source projects. Huge thanks to:
- Ollama โ Local LLM runtime for AI narration and text extraction
- Qwen 2.5 VL by Alibaba Cloud โ Vision-language model for panel understanding
- Ultralytics YOLOv8 โ Object detection framework for panel detection
- Magi v2 by Ragav Sachdeva โ Manga/comic page understanding and speaker-aware transcription
- EasyOCR by JaidedAI โ Text detection for speech bubble cleaning
- Manga OCR by kha-white โ OCR tuned for manga/manhwa text
- Google Gemini API โ Cloud AI narration alternative
- Anthropic Claude API โ Cloud AI narration alternative
- Edge TTS by rany2 โ Free Microsoft Edge TTS (default engine)
- Kokoro TTS by hexgrad โ High-quality local text-to-speech
- ElevenLabs โ Premium cloud TTS
- MoviePy by Zulko โ Python video editing
- FFmpeg โ The backbone of all media processing
- OpenCV โ Computer vision operations
- Pillow โ Image manipulation
- Beautiful Soup by Leonard Richardson โ HTML parsing
- curl_cffi โ Cloudflare-friendly HTTP with TLS fingerprinting
- CloudScraper by VeNoMouS โ Cloudflare challenge bypass
- DrissionPage by g1879 โ Headless browser fallback for protected sites
- Flask โ Python web framework
- Flask-SocketIO โ Real-time progress streaming
- lbpcascade_animeface by nagadomi โ Anime/manga face detection cascade
- PyTorch โ Deep learning framework
- Hugging Face Transformers โ Model loading for Magi v2
- MyAnimeList โ Character data source for the cast system
This project is licensed under the MIT License.
Contributions are welcome! See CONTRIBUTING.md for guidelines.
Made with โค๏ธ for the manhwa community
