Jifeng Dai

21

Papers

2,706

Total Citations

Papers (21)

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

The All-Seeing Project V2: Towards General Relation Comprehension of the Open World

Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and Editing

Point2RBox: Combine Knowledge from Synthetic Visual Patterns for End-to-end Oriented Object Detection with Single Point Supervision

Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

PUMA: Empowering Unified MLLM with Multi-granular Visual Generation

Docopilot: Improving Multimodal Models for Document-Level Understanding

MI-DETR: An Object Detection Model with Multi-time Inquiries Mechanism

MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding

OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis

Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD Drawings

RoboCodeX: Multimodal Code Generation for Robotic Behavior Synthesis

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding

LangBridge: Interpreting Image as a Combination of Language Embeddings

Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft

Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models