Skip to content

Repository files navigation

🧠 Smart-Scraper-RSS

AI-Powered Content Curation for Chinese Social Media

English | 中文


它是什么?

Smart-Scraper-RSS 是一个开源自部署的智能内容聚合系统。它从哔哩哔哩、小红书等中国社交媒体平台抓取内容,通过 LLM(大语言模型)进行质量评分和智能过滤,最终生成干净的 RSS 订阅源。

解决的问题: 中国社交媒体平台充斥着广告、引战、低质量内容。手动刷信息流浪费时间且体验差。

解决方案: 订阅一个 AI 预筛选过的 RSS 源——每条内容都经过质量评分(0-100),低于阈值的自动过滤,你只看到真正有价值的内容。

为什么不用现有工具?

工具 开源自部署 中国平台支持 AI 智能评分 智能过滤
RSSHub (44.7k⭐) ✅ ✅ 强 ❌ ❌ 仅正则
RSS-Bridge ✅ ❌ ❌ ❌
Newscope ✅ ❌ ✅ ✅
BestBlogs.dev ❌ SaaS ✅ ✅ ✅
Feedly ❌ SaaS ❌ ✅ ✅
Smart-Scraper-RSS ✅ ✅ ✅ ✅

核心差异: 目前没有一个开源项目同时支持中国社交媒体平台 + LLM 驱动的智能内容评分和过滤。

RSSHub 能生成 RSS,但过滤只靠正则表达式,没有语义理解。Newscope 有 AI 评分,但不支持中国平台。BestBlogs.dev 最接近,但它是闭源 SaaS。

Smart-Scraper-RSS 填补了这个空白。

核心功能

🕷️ 智能爬虫

  • 多平台支持: B站视频(含CC字幕提取)、小红书笔记、小黑盒资讯、酷安动态
  • 反检测: 浏览器指纹伪装、Cookie 持久化、人类行为模拟
  • 验证码处理: OpenCV 滑块缺口识别 + Bezier 曲线轨迹模拟
  • 增量抓取: URL 去重,只抓取新内容

🤖 AI 内容分析

  • 质量评分: 0-100 分,基于内容深度、原创性、信息价值
  • 广告检测: 自动识别软文、推广、营销内容
  • 风险标记: 检测引战、对立、负面情绪内容
  • 智能摘要: 生成内容摘要,B站视频不看也能懂
  • 情感分析: 正面/中性/负面
  • 支持多模型: DeepSeek、ChatGPT、或任何 OpenAI 兼容 API

📡 RSS 输出

  • 标准 RSS 2.0: 兼容所有主流阅读器(Feedly、Reeder、Inoreader 等)
  • 智能过滤: 只输出高质量内容(可配置最低分数阈值)
  • 丰富条目: 每条 RSS 包含 AI 摘要、评分徽章、风险等级、情感标签
  • 实时更新: 支持定时自动刷新

🖥️ 现代化界面

  • Web 管理面板: 浏览器访问,无需命令行
  • 仪表盘: 实时统计、最近活动、快速操作
  • 数据源管理: 添加/编辑/删除/启用/禁用
  • AI 配置: 模型选择、评分阈值、过滤规则
  • 内容浏览: 查看抓取结果、AI 分析详情

What is it?

Smart-Scraper-RSS is a self-hosted, open-source intelligent content aggregation system. It scrapes content from Chinese social media platforms (Bilibili, Xiaohongshu, etc.), runs it through an LLM for quality scoring and smart filtering, and generates a clean RSS feed.

The problem: Chinese social media platforms are filled with ads, inflammatory posts, and low-quality content. Manually browsing them is a waste of time.

The solution: Subscribe to an AI-curated RSS feed — every item is quality-scored (0-100), and anything below your threshold is automatically filtered out. You only see genuinely valuable content.

Why not use existing tools?

See the comparison table above. No existing open-source project combines Chinese social media support + LLM-powered intelligent content scoring and filtering. Smart-Scraper-RSS fills this gap.

Core Features

🕷️ Smart Scraper

  • Multi-platform: Bilibili (with CC subtitle extraction), Xiaohongshu, Xiaoheihe, CoolAPK
  • Anti-detection: Browser fingerprint masking, cookie persistence, human behavior simulation
  • Captcha handling: OpenCV slider gap detection + Bezier curve trajectory simulation
  • Incremental: URL deduplication, only scrapes new content

🤖 AI Content Analysis

  • Quality scoring: 0-100 based on depth, originality, information value
  • Ad detection: Automatically identifies sponsored/promotional content
  • Risk flagging: Detects inflammatory, polarizing, negative content
  • Smart summaries: Understand a Bilibili video without watching it
  • Sentiment analysis: Positive / Neutral / Negative
  • Multi-model support: DeepSeek, ChatGPT, or any OpenAI-compatible API

📡 RSS Output

  • Standard RSS 2.0: Compatible with all major readers (Feedly, Reeder, Inoreader, etc.)
  • Smart filtering: Only outputs high-quality content (configurable minimum score threshold)
  • Rich entries: Each RSS item includes AI summary, score badge, risk level, sentiment
  • Real-time updates: Scheduled automatic refresh

🖥️ Modern Interface

  • Web dashboard: Browser-based, no command line needed
  • Dashboard: Real-time stats, recent activity, quick actions
  • Source management: Add/edit/delete/enable/disable sources
  • AI configuration: Model selection, scoring thresholds, filter rules
  • Content browser: View scraped results and AI analysis details

Tech Stack

Layer Technology Why
Web Framework FastAPI + Uvicorn High-performance async API, auto-generated docs
Frontend Jinja2 + HTMX + Tailwind CSS Server-side rendering with dynamic updates, no build step
Scraper DrissionPage Chromium-based automation with anti-fingerprinting
AI Engine OpenAI-compatible API Works with DeepSeek, ChatGPT, local LLMs, etc.
Database SQLModel + SQLite Type-safe ORM, zero-config database
Scheduler APScheduler Reliable background job scheduling
RSS feedgen Standard RSS 2.0 generation

Quick Start

Prerequisites

  • Python 3.10+
  • An OpenAI-compatible API key (DeepSeek recommended for cost-effectiveness)

Install

git clone https://github.com/tianxingleo/Smart-Scraper-RSS.git
cd Smart-Scraper-RSS
pip install -r requirements.txt

Configure

cp .env.example .env
# Edit .env with your API key and settings

Key settings in .env:

# AI Configuration
AI_API_KEY=sk-your-api-key
AI_BASE_URL=https://api.deepseek.com/v1
AI_MODEL=deepseek-chat

# RSS Feed
RSS_TITLE=My Curated Feed
RSS_DESCRIPTION=AI-curated content from Chinese social media

# Scraper
HEADLESS=true

Run

python -m app.main

Open http://localhost:8081 in your browser.

Docker (Optional)

docker build -t smart-scraper-rss .
docker run --env-file .env -e WEB_HOST=0.0.0.0 -p 8081:8081 -v ./data:/app/data smart-scraper-rss

Architecture

Smart-Scraper-RSS/
├── app/
│   ├── main.py              # Application entry point
│   ├── config.py             # Configuration management
│   ├── api/                  # FastAPI routes
│   │   ├── routes.py         # API endpoints
│   │   └── deps.py           # Dependency injection
│   ├── scraper/              # Scraper engine
│   │   ├── base.py           # Base scraper class
│   │   ├── browser.py        # Browser manager (DrissionPage)
│   │   ├── strategies/       # Platform-specific scrapers
│   │   │   ├── bilibili.py
│   │   │   ├── xiaohongshu.py
│   │   │   ├── xiaoheihe.py
│   │   │   └── coolapk.py
│   │   └── utils/
│   │       └── captcha.py    # Captcha solving
│   ├── ai/                   # AI analysis engine
│   │   ├── client.py         # LLM API client
│   │   ├── prompts.py        # Prompt templates
│   │   └── analyzer.py       # Content analysis pipeline
│   ├── models/               # Database models
│   │   ├── source.py         # RSS source configuration
│   │   └── item.py           # Scraped content items
│   ├── rss/                  # RSS feed generation
│   │   └── generator.py      # RSS 2.0 generator
│   ├── services/             # Business logic
│   │   ├── scraper_service.py
│   ├── scheduler/            # Job scheduling
│   │   └── manager.py        # APScheduler wrapper
│   └── templates/            # Jinja2 HTML templates
│       ├── base.html
│       ├── dashboard.html
│       ├── sources.html
│       ├── settings.html
│       ├── content.html
│       ├── cookies.html
│       └── watch.html
├── data/                     # Database and browser profiles
├── tests/                    # Test suite
├── requirements.txt
├── Dockerfile
└── .env.example

How It Works

┌─────────────────────────────────────────────────────────────┐
│                     Smart-Scraper-RSS                       │
│                                                             │
│  ┌──────────┐    ┌──────────┐    ┌──────────┐    ┌───────┐ │
│  │ Scheduler │───▶│ Scraper  │───▶│ AI Engine│───▶│  RSS  │ │
│  │ (定时任务) │    │ (爬虫引擎)│    │ (AI分析)  │    │(输出) │ │
│  └──────────┘    └──────────┘    └──────────┘    └───────┘ │
│       │               │               │              │      │
│       ▼               ▼               ▼              ▼      │
│  ┌──────────┐    ┌──────────┐    ┌──────────┐    ┌───────┐ │
│  │ Database │    │ Browser  │    │ LLM API  │    │ Feed  │ │
│  │ (SQLite) │    │(Drission)│    │(DeepSeek)│    │ XML   │ │
│  └──────────┘    └──────────┘    └──────────┘    └───────┘ │
│                                                             │
│  ┌─────────────────────────────────────────────────────┐   │
│  │              Web Dashboard (FastAPI)                 │   │
│  │  ┌─────────┐  ┌─────────┐  ┌─────────┐  ┌────────┐ │   │
│  │  │仪表盘    │  │数据源    │  │设置     │  │内容浏览 │ │   │
│  │  │Dashboard │  │Sources  │  │Settings │  │Content │ │   │
│  │  └─────────┘  └─────────┘  └─────────┘  └────────┘ │   │
│  └─────────────────────────────────────────────────────┘   │
└─────────────────────────────────────────────────────────────┘
         │                                    │
         ▼                                    ▼
  ┌──────────────┐                    ┌──────────────┐
  │  RSS Reader  │                    │   Browser    │
  │(Feedly/Reeder│                    │  (管理面板)   │
  │  /Inoreader) │                    │              │
  └──────────────┘                    └──────────────┘

Configuration

All configuration is via environment variables (.env file):

Variable Default Description
AI_API_KEY - API key for LLM service
AI_BASE_URL https://api.deepseek.com/v1 API endpoint
AI_MODEL deepseek-chat Model name
AI_MIN_SCORE 60 Minimum quality score (0-100) to include in RSS
RSS_TITLE Smart Scraper RSS Feed title
RSS_DESCRIPTION AI-curated content Feed description
HEADLESS true Run browser in headless mode
PROXY - HTTP proxy URL
WEB_PORT 8081 Web dashboard port
DB_PATH data/database.db SQLite database path

Contributing

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

Adding a New Platform

  1. Create a new file in app/scraper/strategies/ (e.g., zhihu.py)
  2. Inherit from BaseScraper and implement the scrape() method
  3. Register the strategy in the scraper service factory
  4. Add the platform to the UI source creation dialog

License

MIT License - see LICENSE for details.


Acknowledgments

  • RSSHub - Inspiration for RSS feed generation
  • Newscope - Inspiration for AI-powered content scoring
  • DrissionPage - Browser automation framework
  • DeepSeek - Cost-effective LLM API

About

idea:爬取互联网信息进行ai判断其价值再利用rss推送

Resources

Contributing

Security policy

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages