Skip to content
#

vilt

Here are 18 public repositories matching this topic...

A benchmark suite for lightweight generative multimodal Vision-Language Models, comparing ViLT and SmolVLM under resource-constrained inference environments. Demonstrates CPU-only deployment, model evaluation, and multimodal reasoning with images and text, highlighting practical GenAI engineering for real-world applications.

  • Updated Jan 29, 2026
  • Python

PyTorch multimodal VQA with ViLT, official-toolkit VQA v2 val 68.42, reproducible training and Gradio demo / 基于 PyTorch 的多模态 VQA,官方工具验证集 68.42,支持可复现训练与 Gradio 演示

  • Updated Sep 2, 2026
  • Python

Add this topic to your repo

To associate your repository with the vilt topic, visit your repo's landing page and select "manage topics."

Learn more