Models
Datasets
Spaces
Docs
Enterprise
Pricing
Log In
Sign Up

Collections

Discover the best community collections!

Collections including paper arxiv:2409.18042

A collection of EMOVA models (https://emova-ollm.github.io/)

Running on Zero

7

EMOVA Online Interactive Demo

🔥

7

Live Interactive demo for EMOVA with Qwen-2.5 backbone
Emova-ollm/emova-qwen-2-5-3b

Text Generation • 4B • Updated Mar 13, 2025 • 5 • 2
Emova-ollm/emova-qwen-2-5-3b-hf

Feature Extraction • 4B • Updated Mar 13, 2025 • 14 • 5
Emova-ollm/emova-qwen-2-5-7b

Text Generation • 8B • Updated Mar 13, 2025 • 5 • 1

CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation

Paper • 2410.23090 • Published Oct 30, 2024 • 55
SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Paper • 2410.13276 • Published Oct 17, 2024 • 29
Personalized Visual Instruction Tuning

Paper • 2410.07113 • Published Oct 9, 2024 • 70
Differential Transformer

Paper • 2410.05258 • Published Oct 7, 2024 • 180

Multimodal LLMs

Building and better understanding vision-language models: insights and future directions

Paper • 2408.12637 • Published Aug 22, 2024 • 133
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Paper • 2408.11039 • Published Aug 20, 2024 • 63
Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Paper • 2408.16725 • Published Aug 29, 2024 • 53
Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Paper • 2408.15998 • Published Aug 28, 2024 • 86

iVideoGPT: Interactive VideoGPTs are Scalable World Models

Paper • 2405.15223 • Published May 24, 2024 • 17
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models

Paper • 2405.15574 • Published May 24, 2024 • 55
An Introduction to Vision-Language Modeling

Paper • 2405.17247 • Published May 27, 2024 • 90
Matryoshka Multimodal Models

Paper • 2405.17430 • Published May 27, 2024 • 34

A collection of EMOVA datasets (https://emova-ollm.github.io/)

Emova-ollm/emova-alignment-7m

Viewer • Updated Mar 14, 2025 • 6.18M • 5.77k • 5
Emova-ollm/emova-sft-4m

Viewer • Updated Mar 14, 2025 • 4.31M • 1.09k • 5
Emova-ollm/emova-sft-speech-231k

Viewer • Updated Mar 14, 2025 • 231k • 55 • 3
Emova-ollm/emova-sft-speech-eval

Viewer • Updated Mar 14, 2025 • 3.76k • 23 • 1

impira/layoutlm-document-qa

Document Question Answering • 0.1B • Updated Mar 18, 2023 • 7.7k • 1.15k
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Paper • 2409.18042 • Published Sep 26, 2024 • 39
LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects

Paper • 2504.19838 • Published Apr 28, 2025 • 22

Papers I want to read

Papers in my to-read list

RLHF Workflow: From Reward Modeling to Online RLHF

Paper • 2405.07863 • Published May 13, 2024 • 71
Chameleon: Mixed-Modal Early-Fusion Foundation Models

Paper • 2405.09818 • Published May 16, 2024 • 132
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models

Paper • 2405.15574 • Published May 24, 2024 • 55
An Introduction to Vision-Language Modeling

Paper • 2405.17247 • Published May 27, 2024 • 90

A collection of EMOVA models (https://emova-ollm.github.io/)

Running on Zero

7

EMOVA Online Interactive Demo

🔥

7

Live Interactive demo for EMOVA with Qwen-2.5 backbone
Emova-ollm/emova-qwen-2-5-3b

Text Generation • 4B • Updated Mar 13, 2025 • 5 • 2
Emova-ollm/emova-qwen-2-5-3b-hf

Feature Extraction • 4B • Updated Mar 13, 2025 • 14 • 5
Emova-ollm/emova-qwen-2-5-7b

Text Generation • 8B • Updated Mar 13, 2025 • 5 • 1

A collection of EMOVA datasets (https://emova-ollm.github.io/)

Emova-ollm/emova-alignment-7m

Viewer • Updated Mar 14, 2025 • 6.18M • 5.77k • 5
Emova-ollm/emova-sft-4m

Viewer • Updated Mar 14, 2025 • 4.31M • 1.09k • 5
Emova-ollm/emova-sft-speech-231k

Viewer • Updated Mar 14, 2025 • 231k • 55 • 3
Emova-ollm/emova-sft-speech-eval

Viewer • Updated Mar 14, 2025 • 3.76k • 23 • 1

CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation

Paper • 2410.23090 • Published Oct 30, 2024 • 55
SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Paper • 2410.13276 • Published Oct 17, 2024 • 29
Personalized Visual Instruction Tuning

Paper • 2410.07113 • Published Oct 9, 2024 • 70
Differential Transformer

Paper • 2410.05258 • Published Oct 7, 2024 • 180

impira/layoutlm-document-qa

Document Question Answering • 0.1B • Updated Mar 18, 2023 • 7.7k • 1.15k
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Paper • 2409.18042 • Published Sep 26, 2024 • 39
LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects

Paper • 2504.19838 • Published Apr 28, 2025 • 22

Multimodal LLMs

Building and better understanding vision-language models: insights and future directions

Paper • 2408.12637 • Published Aug 22, 2024 • 133
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Paper • 2408.11039 • Published Aug 20, 2024 • 63
Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Paper • 2408.16725 • Published Aug 29, 2024 • 53
Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Paper • 2408.15998 • Published Aug 28, 2024 • 86

Papers I want to read

Papers in my to-read list

RLHF Workflow: From Reward Modeling to Online RLHF

Paper • 2405.07863 • Published May 13, 2024 • 71
Chameleon: Mixed-Modal Early-Fusion Foundation Models

Paper • 2405.09818 • Published May 16, 2024 • 132
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models

Paper • 2405.15574 • Published May 24, 2024 • 55
An Introduction to Vision-Language Modeling

Paper • 2405.17247 • Published May 27, 2024 • 90

iVideoGPT: Interactive VideoGPTs are Scalable World Models

Paper • 2405.15223 • Published May 24, 2024 • 17
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models

Paper • 2405.15574 • Published May 24, 2024 • 55
An Introduction to Vision-Language Modeling

Paper • 2405.17247 • Published May 27, 2024 • 90
Matryoshka Multimodal Models

Paper • 2405.17430 • Published May 27, 2024 • 34

Company

TOS Privacy About Careers

Website

Models Datasets Spaces Pricing Docs