Skip to content
#

vision-foundation-model

Here are 24 public repositories matching this topic...

A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs). This repository aggregates surveys, blog posts, and research papers that explore how LMMs represent, transform, and align multimodal information internally.

  • Updated Mar 4, 2026

Official implementation of VGGT-DP: Generalizable Robot Control via Vision Foundation Models (arXiv:2509.18778). A diffusion policy that uses the frozen VGGT vision foundation model as a geometry-aware visual encoder for robot manipulation, evaluated on MetaWorld.

  • Updated Sep 6, 2026
  • Python

Add this topic to your repo

To associate your repository with the vision-foundation-model topic, visit your repo's landing page and select "manage topics."

Learn more