Skip to content

Latest commit

ย 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

AniFlow Logo

AniFlow

Turn manhwa & webtoon chapters into narrated recap videos โ€” fully automated.

Quick Start Python License Platform


AniFlow downloads manhwa/webtoon chapters, detects and crops panels using YOLO, extracts speech-bubble text, generates AI narration, converts it to speech, and renders video clips โ€” all through a browser-based UI.

โœจ Features

  • ๐ŸŒ Web Scraper โ€” download chapters from manhuaus, manhwatop, MangaDex, and more
  • ๐Ÿ” YOLO Panel Detection โ€” custom-trained models detect panels and speech bubbles
  • ๐Ÿงน Text Removal โ€” automatically clean speech bubbles from art panels
  • ๐Ÿค– AI Narration โ€” local (Ollama) or cloud (Gemini / Claude) narration
  • ๐ŸŽ™๏ธ Multi-TTS โ€” Edge TTS (free), Kokoro (local), or ElevenLabs (cloud)
  • ๐ŸŽฌ Video Rendering โ€” Ken Burns animation, blur backgrounds, BGM, watermarks
  • ๐Ÿ‘ฅ Character Casting โ€” import from MyAnimeList or scan faces from chapters
  • ๐Ÿ“ฆ Batch Mode โ€” process dozens of chapters unattended

Two Pipelines

Dialogue Recap Narrated Recap
Panels Original panels with speech bubbles Clean art panels (bubbles removed)
Audio Extracted dialogue read aloud AI-written YouTube-style recap narration
Best for Letting the original dialogue shine YouTube story recaps

๐Ÿ“‚ See it in action! Check the example_output/ folder for sample clips from both modes.

๐ŸŽฅ Video Tutorials

Tutorial Link
๐Ÿ’ฌ Dialogue Recap โ€” Single Chapter Watch on YouTube
๐Ÿ’ฌ Dialogue Recap โ€” Batch Mode (50+ Chapters) Watch on YouTube
๐ŸŽ™๏ธ Narrated Recap โ€” Automatic Pipeline Watch on YouTube
๐ŸŽ™๏ธ Narrated Recap โ€” Batch Mode (Full Series) Watch on YouTube

๐Ÿ“‹ Requirements

Dependency Purpose Install
Python 3.10 โ€“ 3.12 Runtime python.org
FFmpeg Video/audio processing brew install ffmpeg (macOS) ยท sudo apt install ffmpeg (Linux) ยท ffmpeg.org (Windows)
9router / Ollama Local AI narration & text extraction 9router / ollama.com
~10 GB disk Python env + model weights โ€”

GPU recommended but not required. Apple Silicon Macs and NVIDIA GPUs will significantly speed up panel detection and Kokoro TTS.


๐Ÿš€ Quick Start

Option A: One-Command Setup (Recommended)

git clone https://github.com/aashish254/Aniflow.git
cd AniFlow
chmod +x setup.sh
./setup.sh

The setup script will check dependencies, create a virtual environment, install packages, pull Ollama models, and start the app.

Option B: Manual Setup

# 1. Clone the repository
git clone https://github.com/aashish254/Aniflow.git
cd AniFlow

# 2. Create a virtual environment (Python 3.12 recommended)
py -3.12 -m venv venv             # Windows
# python3 -m venv venv            # macOS / Linux

# 3. Activate virtual environment
# Windows (PowerShell):
.\venv\Scripts\Activate.ps1
# macOS / Linux:
# source venv/bin/activate

# 4. Install Python dependencies (~10-20 min, torch is large)
pip install --upgrade pip
pip install -r requirements.txt

# 5. Configure 9router (Alternative to Ollama)
# If you use 9router, set these environment variables before running:
# Windows (PowerShell):
# $env:USE_NINE_ROUTER="1"
# $env:NINE_ROUTER_HOST="http://127.0.0.1:20128"
# $env:VISION_MODEL="Free-Tier"
# $env:WRITER_MODEL="Free-Tier"
# $env:TEXT_WRITER_MODEL="Free-Tier"

# 6. Pull Ollama models (Skip if using 9router)
# ollama pull qwen2.5vl:3b
# ollama pull qwen2.5:14b

# 7. Run the app
python app.py

# 8. Open UI in browser:
# http://localhost:8080

๐Ÿ‡ฎ๐Ÿ‡ฉ Panduan Bahasa Indonesia

๐Ÿ“‹ Persyaratan

Dependensi Fungsi Instalasi
Python 3.10 โ€“ 3.12 Runtime python.org
FFmpeg Pemrosesan video/audio brew install ffmpeg (macOS) ยท sudo apt install ffmpeg (Linux) ยท ffmpeg.org (Windows)
9router / Ollama AI lokal untuk narasi & ekstraksi teks 9router / ollama.com
~10 GB disk Python env + model weights โ€”

GPU disarankan tapi tidak wajib. Apple Silicon Mac dan GPU NVIDIA akan sangat mempercepat deteksi panel dan Kokoro TTS.

๐Ÿš€ Cara Menjalankan Aplikasi (Windows & 9router)

Langkah 1: Pastikan 9router Aktif

Pastikan 9router sudah berjalan di latar belakang (port default: http://127.0.0.1:20128).

Langkah 2: Jalankan di PowerShell

# 1. Masuk folder proyek
cd D:\Aniflow

# 2. Aktifkan virtual environment
.\venv\Scripts\Activate.ps1

# 3. Set environment variables untuk 9router
$env:USE_NINE_ROUTER="1"
$env:NINE_ROUTER_HOST="http://127.0.0.1:20128"
$env:VISION_MODEL="Free-Tier"
$env:WRITER_MODEL="Free-Tier"
$env:TEXT_WRITER_MODEL="Free-Tier"

# (Opsional) Jika 9router butuh API Key
# $env:NINE_ROUTER_API_KEY="API_KEY_ANDA"

# 4. Jalankan server AniFlow
python app.py

Langkah 3: Buka Browser

Akses antarmuka web di:

http://localhost:8080

โš™๏ธ Configuration

setup.bat


Or follow the manual steps above using `venv\Scripts\activate` instead.

---

## ๐ŸŽฌ How to Use

### Narrated Recap (Recommended)

1. **Source** โ€” Paste a manhwa URL โ†’ **Analyze** โ†’ pick chapters โ†’ **Download**
   (or point to a local folder of images)
2. **Cast** *(optional)* โ€” Import characters from MyAnimeList, or scan chapter faces
   and assign names. The narrator uses real names ("Jin-Wooโ€ฆ") instead of "the dark-haired guy."
3. **Stitch** โ€” Stitch โ†’ Detect (bubble panels) โ†’ Crop โ†’ Stitch โ†’ Detect (art panels) โ†’ Crop.
   Boxes are editable โ€” drag, resize, or draw new ones.
4. **Review** โ€” Check cropped panels; delete junk or re-crop.
5. **Extract** โ€” AI reads speech-bubble text.
6. **Narrate** โ€” Choose your AI backend:
   - ๐Ÿฆ™ **Ollama (local)** โ€” free & offline. `qwen2.5vl:7b` (vision) or `qwen2.5:14b` (text-only).
   - โœจ **Gemini API** โ€” paste multiple API keys (auto-rotation on quota). Free keys at [aistudio.google.com](https://aistudio.google.com).
   - ๐Ÿค– **Claude API** โ€” single key.
7. **Review Text** โ€” Edit any narration line or re-narrate individual panels.
8. **Audio** โ€” TTS with Edge (free), Kokoro (local), or ElevenLabs.
9. **Video** โ€” Renders clips with Ken Burns animation, blur backgrounds, BGM.
10. **Review Clips / Sort Clips** โ€” Preview, download, or batch-merge.

### Dialogue Recap

Same flow minus the cast/narrate steps โ€” extracted dialogue becomes the audio,
panels keep their original speech bubbles.

### Batch Mode

Download chapters 1โ€“68, enable **Batch Mode**, hit **Run Full Batch** โ€” every step
runs across all chapters automatically. Progress is streamed live.

---

## ๐Ÿ“ Project Structure

AniFlow/ โ”œโ”€โ”€ app.py # Flask web app + API routes โ”œโ”€โ”€ config.py # All configuration & path settings โ”œโ”€โ”€ requirements.txt # Python dependencies โ”œโ”€โ”€ setup.sh / setup.bat # One-command setup scripts โ”œโ”€โ”€ start.sh # Quick-start script (macOS/Linux) โ”‚ โ”œโ”€โ”€ pipeline/ # Core processing modules โ”‚ โ”œโ”€โ”€ web_scraper.py # Chapter downloading (multi-site) โ”‚ โ”œโ”€โ”€ stitcher.py # Image stitching โ”‚ โ”œโ”€โ”€ panel_extractor_yolo.py # YOLO panel detection โ”‚ โ”œโ”€โ”€ cropper.py # Panel cropping โ”‚ โ”œโ”€โ”€ text_extractor.py # OCR / text extraction โ”‚ โ”œโ”€โ”€ text_remover.py # Speech bubble cleaning โ”‚ โ”œโ”€โ”€ narrator.py # Local AI narration (Ollama) โ”‚ โ”œโ”€โ”€ gemini_narrator.py # Cloud narration (Gemini API) โ”‚ โ”œโ”€โ”€ scene_narrator.py # Scene-level narration โ”‚ โ”œโ”€โ”€ tts_engine.py # Edge TTS engine โ”‚ โ”œโ”€โ”€ kokoro_tts.py # Kokoro local TTS โ”‚ โ”œโ”€โ”€ elevenlabs_tts.py # ElevenLabs cloud TTS โ”‚ โ”œโ”€โ”€ compositor.py # Video rendering โ”‚ โ”œโ”€โ”€ orchestrator.py # Pipeline orchestration โ”‚ โ”œโ”€โ”€ cast_system.py # Character management โ”‚ โ”œโ”€โ”€ character_builder.py # Face detection & grouping โ”‚ โ”œโ”€โ”€ mal_character_importer.py # MAL character import โ”‚ โ”œโ”€โ”€ script_rewriter.py # Narration style rewriter โ”‚ โ”œโ”€โ”€ capcut_director.py # CapCut export โ”‚ โ””โ”€โ”€ panel_filter.py # Panel quality filtering โ”‚ โ”œโ”€โ”€ models/ # YOLO model weights (committed) โ”‚ โ”œโ”€โ”€ manhwa_panel_cropped.pt # Art panel detector โ”‚ โ””โ”€โ”€ panel_with_bubble_text.pt # Bubble panel detector โ”‚ โ”œโ”€โ”€ templates/ # HTML templates โ”œโ”€โ”€ static/ # CSS, JS, logos โ”œโ”€โ”€ assets/ # BGM, watermarks, intros, outros โ””โ”€โ”€ tests/ # Test & diagnostic scripts


### Generated Folders (gitignored)

| Folder | Contents |
|---|---|
| `manhwa_download/` | Downloaded chapter images + per-chapter state |
| `manhwa_stitch/` | Stitched full-page strips |
| `manhwa_detect/` | Detection metadata (bounding boxes) |
| `manhwa_crop/` | Cropped individual panels |
| `manhwa_text/` | Extracted dialogue JSON |
| `manhwa_audio/` | Generated TTS audio files |
| `manhwa_clips/` | Final rendered video clips |
| `output/` | Legacy output directory |
| `~/RECAP/Cast Data/` | Your character casts (faces + profiles) |

All generated data is **regenerable** โ€” delete any folder and re-run the pipeline.

---

## โš™๏ธ Configuration

### API Keys (Optional)

Copy `.env.example` to `.env` and fill in keys only for services you want to use:

```bash
cp .env.example .env
Key Service Required?
GEMINI_API_KEY Google Gemini narration Only if using Gemini
ANTHROPIC_API_KEY Claude narration Only if using Claude
ELEVENLABS_API_KEY ElevenLabs TTS Only if using ElevenLabs
MAL_CLIENT_ID MyAnimeList import Only if importing characters from MAL

Ollama + Edge TTS work completely offline โ€” no API keys needed for the core pipeline.

Ollama Models

Model Role Size Required?
qwen2.5vl:3b Vision parser (reads panels) ~2 GB โœ… Required
qwen2.5:14b Text-only narrator ~9 GB ๐Ÿ“Œ Recommended
qwen2.5vl:7b Vision narrator (sees panels + cast) ~5 GB Optional

๐Ÿ”ง Troubleshooting

Problem Solution
"Address already in use" Another instance is running: lsof -ti :8080 | xargs kill -9
Ollama dot is red Start Ollama and pull models: ollama pull qwen2.5vl:3b
Downloads hang Protected site โ€” the app auto-switches to headless browser (~40s). MangaDex URLs are most reliable.
Bubbles not removed EasyOCR inpainter handles clear bubbles; complex ones can be fixed with the crop tool in Review.
ModuleNotFoundError Make sure your venv is activated: source venv/bin/activate
Torch install fails Try: pip install torch torchvision --index-url https://download.pytorch.org/whl/cpu

๐Ÿ™ Acknowledgments & Credits

AniFlow is built on the shoulders of many incredible open-source projects. Huge thanks to:

AI & Machine Learning

  • Ollama โ€” Local LLM runtime for AI narration and text extraction
  • Qwen 2.5 VL by Alibaba Cloud โ€” Vision-language model for panel understanding
  • Ultralytics YOLOv8 โ€” Object detection framework for panel detection
  • Magi v2 by Ragav Sachdeva โ€” Manga/comic page understanding and speaker-aware transcription
  • EasyOCR by JaidedAI โ€” Text detection for speech bubble cleaning
  • Manga OCR by kha-white โ€” OCR tuned for manga/manhwa text
  • Google Gemini API โ€” Cloud AI narration alternative
  • Anthropic Claude API โ€” Cloud AI narration alternative

Text-to-Speech

  • Edge TTS by rany2 โ€” Free Microsoft Edge TTS (default engine)
  • Kokoro TTS by hexgrad โ€” High-quality local text-to-speech
  • ElevenLabs โ€” Premium cloud TTS

Video & Image Processing

  • MoviePy by Zulko โ€” Python video editing
  • FFmpeg โ€” The backbone of all media processing
  • OpenCV โ€” Computer vision operations
  • Pillow โ€” Image manipulation

Web Scraping

  • Beautiful Soup by Leonard Richardson โ€” HTML parsing
  • curl_cffi โ€” Cloudflare-friendly HTTP with TLS fingerprinting
  • CloudScraper by VeNoMouS โ€” Cloudflare challenge bypass
  • DrissionPage by g1879 โ€” Headless browser fallback for protected sites

Web Framework

Other


๐Ÿ“œ License

This project is licensed under the MIT License.


๐Ÿค Contributing

Contributions are welcome! See CONTRIBUTING.md for guidelines.


Made with โค๏ธ for the manhwa community

About

๐ŸŽฌ Autonomous AI Studio for Manhwa & Webtoon Recaps โ€” Vision Pipeline (YOLOv8 + Magi v2), Multi-Voice TTS & Video Synthesis

Topics

Resources

Contributing

Security policy

Stars

14 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages