AI-Powered Content Curation for Chinese Social Media
Smart-Scraper-RSS 是一个开源自部署的智能内容聚合系统。它从哔哩哔哩、小红书等中国社交媒体平台抓取内容,通过 LLM(大语言模型)进行质量评分和智能过滤,最终生成干净的 RSS 订阅源。
解决的问题: 中国社交媒体平台充斥着广告、引战、低质量内容。手动刷信息流浪费时间且体验差。
解决方案: 订阅一个 AI 预筛选过的 RSS 源——每条内容都经过质量评分(0-100),低于阈值的自动过滤,你只看到真正有价值的内容。
| 工具 | 开源自部署 | 中国平台支持 | AI 智能评分 | 智能过滤 |
|---|---|---|---|---|
| RSSHub (44.7k⭐) | ✅ | ✅ 强 | ❌ | ❌ 仅正则 |
| RSS-Bridge | ✅ | ❌ | ❌ | ❌ |
| Newscope | ✅ | ❌ | ✅ | ✅ |
| BestBlogs.dev | ❌ SaaS | ✅ | ✅ | ✅ |
| Feedly | ❌ SaaS | ❌ | ✅ | ✅ |
| Smart-Scraper-RSS | ✅ | ✅ | ✅ | ✅ |
核心差异: 目前没有一个开源项目同时支持中国社交媒体平台 + LLM 驱动的智能内容评分和过滤。
RSSHub 能生成 RSS,但过滤只靠正则表达式,没有语义理解。Newscope 有 AI 评分,但不支持中国平台。BestBlogs.dev 最接近,但它是闭源 SaaS。
Smart-Scraper-RSS 填补了这个空白。
- 多平台支持: B站视频(含CC字幕提取)、小红书笔记、小黑盒资讯、酷安动态
- 反检测: 浏览器指纹伪装、Cookie 持久化、人类行为模拟
- 验证码处理: OpenCV 滑块缺口识别 + Bezier 曲线轨迹模拟
- 增量抓取: URL 去重,只抓取新内容
- 质量评分: 0-100 分,基于内容深度、原创性、信息价值
- 广告检测: 自动识别软文、推广、营销内容
- 风险标记: 检测引战、对立、负面情绪内容
- 智能摘要: 生成内容摘要,B站视频不看也能懂
- 情感分析: 正面/中性/负面
- 支持多模型: DeepSeek、ChatGPT、或任何 OpenAI 兼容 API
- 标准 RSS 2.0: 兼容所有主流阅读器(Feedly、Reeder、Inoreader 等)
- 智能过滤: 只输出高质量内容(可配置最低分数阈值)
- 丰富条目: 每条 RSS 包含 AI 摘要、评分徽章、风险等级、情感标签
- 实时更新: 支持定时自动刷新
- Web 管理面板: 浏览器访问,无需命令行
- 仪表盘: 实时统计、最近活动、快速操作
- 数据源管理: 添加/编辑/删除/启用/禁用
- AI 配置: 模型选择、评分阈值、过滤规则
- 内容浏览: 查看抓取结果、AI 分析详情
Smart-Scraper-RSS is a self-hosted, open-source intelligent content aggregation system. It scrapes content from Chinese social media platforms (Bilibili, Xiaohongshu, etc.), runs it through an LLM for quality scoring and smart filtering, and generates a clean RSS feed.
The problem: Chinese social media platforms are filled with ads, inflammatory posts, and low-quality content. Manually browsing them is a waste of time.
The solution: Subscribe to an AI-curated RSS feed — every item is quality-scored (0-100), and anything below your threshold is automatically filtered out. You only see genuinely valuable content.
See the comparison table above. No existing open-source project combines Chinese social media support + LLM-powered intelligent content scoring and filtering. Smart-Scraper-RSS fills this gap.
- Multi-platform: Bilibili (with CC subtitle extraction), Xiaohongshu, Xiaoheihe, CoolAPK
- Anti-detection: Browser fingerprint masking, cookie persistence, human behavior simulation
- Captcha handling: OpenCV slider gap detection + Bezier curve trajectory simulation
- Incremental: URL deduplication, only scrapes new content
- Quality scoring: 0-100 based on depth, originality, information value
- Ad detection: Automatically identifies sponsored/promotional content
- Risk flagging: Detects inflammatory, polarizing, negative content
- Smart summaries: Understand a Bilibili video without watching it
- Sentiment analysis: Positive / Neutral / Negative
- Multi-model support: DeepSeek, ChatGPT, or any OpenAI-compatible API
- Standard RSS 2.0: Compatible with all major readers (Feedly, Reeder, Inoreader, etc.)
- Smart filtering: Only outputs high-quality content (configurable minimum score threshold)
- Rich entries: Each RSS item includes AI summary, score badge, risk level, sentiment
- Real-time updates: Scheduled automatic refresh
- Web dashboard: Browser-based, no command line needed
- Dashboard: Real-time stats, recent activity, quick actions
- Source management: Add/edit/delete/enable/disable sources
- AI configuration: Model selection, scoring thresholds, filter rules
- Content browser: View scraped results and AI analysis details
| Layer | Technology | Why |
|---|---|---|
| Web Framework | FastAPI + Uvicorn | High-performance async API, auto-generated docs |
| Frontend | Jinja2 + HTMX + Tailwind CSS | Server-side rendering with dynamic updates, no build step |
| Scraper | DrissionPage | Chromium-based automation with anti-fingerprinting |
| AI Engine | OpenAI-compatible API | Works with DeepSeek, ChatGPT, local LLMs, etc. |
| Database | SQLModel + SQLite | Type-safe ORM, zero-config database |
| Scheduler | APScheduler | Reliable background job scheduling |
| RSS | feedgen | Standard RSS 2.0 generation |
- Python 3.10+
- An OpenAI-compatible API key (DeepSeek recommended for cost-effectiveness)
git clone https://github.com/tianxingleo/Smart-Scraper-RSS.git
cd Smart-Scraper-RSS
pip install -r requirements.txtcp .env.example .env
# Edit .env with your API key and settingsKey settings in .env:
# AI Configuration
AI_API_KEY=sk-your-api-key
AI_BASE_URL=https://api.deepseek.com/v1
AI_MODEL=deepseek-chat
# RSS Feed
RSS_TITLE=My Curated Feed
RSS_DESCRIPTION=AI-curated content from Chinese social media
# Scraper
HEADLESS=truepython -m app.mainOpen http://localhost:8081 in your browser.
docker build -t smart-scraper-rss .
docker run --env-file .env -e WEB_HOST=0.0.0.0 -p 8081:8081 -v ./data:/app/data smart-scraper-rssSmart-Scraper-RSS/
├── app/
│ ├── main.py # Application entry point
│ ├── config.py # Configuration management
│ ├── api/ # FastAPI routes
│ │ ├── routes.py # API endpoints
│ │ └── deps.py # Dependency injection
│ ├── scraper/ # Scraper engine
│ │ ├── base.py # Base scraper class
│ │ ├── browser.py # Browser manager (DrissionPage)
│ │ ├── strategies/ # Platform-specific scrapers
│ │ │ ├── bilibili.py
│ │ │ ├── xiaohongshu.py
│ │ │ ├── xiaoheihe.py
│ │ │ └── coolapk.py
│ │ └── utils/
│ │ └── captcha.py # Captcha solving
│ ├── ai/ # AI analysis engine
│ │ ├── client.py # LLM API client
│ │ ├── prompts.py # Prompt templates
│ │ └── analyzer.py # Content analysis pipeline
│ ├── models/ # Database models
│ │ ├── source.py # RSS source configuration
│ │ └── item.py # Scraped content items
│ ├── rss/ # RSS feed generation
│ │ └── generator.py # RSS 2.0 generator
│ ├── services/ # Business logic
│ │ ├── scraper_service.py
│ ├── scheduler/ # Job scheduling
│ │ └── manager.py # APScheduler wrapper
│ └── templates/ # Jinja2 HTML templates
│ ├── base.html
│ ├── dashboard.html
│ ├── sources.html
│ ├── settings.html
│ ├── content.html
│ ├── cookies.html
│ └── watch.html
├── data/ # Database and browser profiles
├── tests/ # Test suite
├── requirements.txt
├── Dockerfile
└── .env.example
┌─────────────────────────────────────────────────────────────┐
│ Smart-Scraper-RSS │
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌───────┐ │
│ │ Scheduler │───▶│ Scraper │───▶│ AI Engine│───▶│ RSS │ │
│ │ (定时任务) │ │ (爬虫引擎)│ │ (AI分析) │ │(输出) │ │
│ └──────────┘ └──────────┘ └──────────┘ └───────┘ │
│ │ │ │ │ │
│ ▼ ▼ ▼ ▼ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌───────┐ │
│ │ Database │ │ Browser │ │ LLM API │ │ Feed │ │
│ │ (SQLite) │ │(Drission)│ │(DeepSeek)│ │ XML │ │
│ └──────────┘ └──────────┘ └──────────┘ └───────┘ │
│ │
│ ┌─────────────────────────────────────────────────────┐ │
│ │ Web Dashboard (FastAPI) │ │
│ │ ┌─────────┐ ┌─────────┐ ┌─────────┐ ┌────────┐ │ │
│ │ │仪表盘 │ │数据源 │ │设置 │ │内容浏览 │ │ │
│ │ │Dashboard │ │Sources │ │Settings │ │Content │ │ │
│ │ └─────────┘ └─────────┘ └─────────┘ └────────┘ │ │
│ └─────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘
│ │
▼ ▼
┌──────────────┐ ┌──────────────┐
│ RSS Reader │ │ Browser │
│(Feedly/Reeder│ │ (管理面板) │
│ /Inoreader) │ │ │
└──────────────┘ └──────────────┘
All configuration is via environment variables (.env file):
| Variable | Default | Description |
|---|---|---|
AI_API_KEY |
- | API key for LLM service |
AI_BASE_URL |
https://api.deepseek.com/v1 |
API endpoint |
AI_MODEL |
deepseek-chat |
Model name |
AI_MIN_SCORE |
60 |
Minimum quality score (0-100) to include in RSS |
RSS_TITLE |
Smart Scraper RSS |
Feed title |
RSS_DESCRIPTION |
AI-curated content |
Feed description |
HEADLESS |
true |
Run browser in headless mode |
PROXY |
- | HTTP proxy URL |
WEB_PORT |
8081 |
Web dashboard port |
DB_PATH |
data/database.db |
SQLite database path |
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
- Create a new file in
app/scraper/strategies/(e.g.,zhihu.py) - Inherit from
BaseScraperand implement thescrape()method - Register the strategy in the scraper service factory
- Add the platform to the UI source creation dialog
MIT License - see LICENSE for details.
- RSSHub - Inspiration for RSS feed generation
- Newscope - Inspiration for AI-powered content scoring
- DrissionPage - Browser automation framework
- DeepSeek - Cost-effective LLM API