AI Explorer · Software Engineering Undergraduate at Tongji University
I study how AI systems listen, look, reason, and revise - and I build tools that make those decisions easier to inspect.
- Multimodal perception: grounding objects and events across audio, vision, motion, and language.
- Reasoning behavior: understanding when longer reasoning helps and when it creates drift.
- AI for research: building inspectable workflows for reading, experimentation, and scientific communication.
Audio-visual segmentation research that uses spectral-kinematic alignment and motion-guided queries to reduce false positives from visually salient but silent objects.
A reproducible pipeline for studying zero, short, and long reasoning budgets before multimodal grounding and segmentation.
A multi-agent paper-to-poster system for controllable design diversity and editable, print-ready PowerPoint outputs. Includes a 621-paper benchmark, 24 templates, and bounded visual-quality repair. Project page · Paper
An auditable NBA draft prediction agent that fuses talent, expert mocks, and market signals across 30 GM personas and 1,500 Monte Carlo scenarios. Built for the AWS Summit Shanghai 2026 hackathon; placed third and advanced to the Macau round.
A Cocos2d-x systems project covering map interaction, character control, collision detection, inventory, and farming simulation mechanics.
Looking for AI research and engineering internships around multimodal learning, LLM reasoning, research agents, and evaluation-heavy systems.
Python · PyTorch · C++ · TypeScript · Computer Vision · Multimodal Learning · Research Tooling
I also enjoy playful systems, game mechanics, and projects that make difficult ideas visible.
