Code and implementation guidelines for the paper ✨Counting Anything. Project Page: https://mengqi-lei.github.io/count-anything-projectpage/
-
Updated
Aug 12, 2026 - Python
Code and implementation guidelines for the paper ✨Counting Anything. Project Page: https://mengqi-lei.github.io/count-anything-projectpage/
OpenVision (ICCV 2025), OpenVision 2 (CVPR 2026), and OpenVision 3
A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs). This repository aggregates surveys, blog posts, and research papers that explore how LMMs represent, transform, and align multimodal information internally.
[--branch main] Face Security Foundation Model via Self-Supervised Facial Representation Learning (CVPR 2025). [--branch FSVFM-extension-R1] FS-VFM extension with scalable face-security visual backbones, Linear Probing, FS-Adapter, and DF40 evaluation support.
Recognize Any Regions
[ICLR 2026] FSOD-VFM: Few-Shot Object Detection with Vision Foundation Models and Graph Diffusion
A Vision Foundation Model for Cine Cardiac Magnetic Resonance Imaging
One-Shot Open Affordance Learning with Foundation Models (CVPR 2024)
[WAICA-26 Best Student Paper] Official repository of "Enhancing Vision Foundation Models via Multimodal Continual Pre-Training"
MonoDINO-DETR: Depth-Enhanced Monocular 3D Object Detection Using a Vision Foundation Model
This repo collects some latest research work of Generative AI. It provides simple implementations to understand the ideas and some follow-up discussions to inspire future work.
[Nature Communications] Official Code for Segment Any Tumour - SAT3D
Implementation of CAST: Contrastive Adaptation and Distillation for Semi-Supervised Instance Segmentation.
"Boosting Gaze Object Prediction via Pixel-level Supervision from Vision Foundation Model"
A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models
[CVPR 2026 Findings] RecycleLoRA: Rank-Revealing QR-Based Dual-LoRA Subspace Adaptation for Domain Generalized Semantic Segmentation
Codebase for probing VFMs and Feature Upsamplers using Intractive Segmentation.
Simple Gradio application integrated with Hugging Face Multimodals to support visual question answering chatbot and more features
Official implementation of VGGT-DP: Generalizable Robot Control via Vision Foundation Models (arXiv:2509.18778). A diffusion policy that uses the frozen VGGT vision foundation model as a geometry-aware visual encoder for robot manipulation, evaluated on MetaWorld.
To associate your repository with the vision-foundation-model topic, visit your repo's landing page and select "manage topics."