Computer Vision and Pattern Recognition

Near-perfect photo-ID of the Hula painted frog with zero-shot deep local-feature matching

Near-perfect photo-ID of the Hula painted frog...

Computer Vision and Pattern Recognition
Avatar
yoavram
68 views
Multilayer Graph Approach to Deep Subspace Clustering

Multilayer Graph Approach to Deep Subspace Clu...

Computer Vision and Pattern Recognition
Avatar
lovro-sindicic
73 views
Label-independent hyperparameter-free self-supervised single-view deep subspace clustering

Label-independent hyperparameter-free self-sup...

Computer Vision and Pattern Recognition
Avatar
lovro-sindicic
80 views
PersonaLive! Expressive Portrait Image Animation for Live Streaming

PersonaLive! Expressive Portrait Image Animati...

Computer Vision and Pattern Recognition
Avatar
Grisha Samokhin
90 views
Mull-Tokens: Modality-Agnostic Latent Thinking

Mull-Tokens: Modality-Agnostic Latent Thinking

Computer Vision and Pattern Recognition
Avatar
librarian
105 views
Linear Gaussian Bounding Box Representation and Ring-Shaped Rotated Convolution for Oriented Object Detection

Linear Gaussian Bounding Box Representation an...

Computer Vision and Pattern Recognition
Avatar
rahulraj Kk
83 views
Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer
  Memory

Point3R: Streaming 3D Reconstruction with Expl...

Computer Vision and Pattern Recognition
Avatar
librarian
391 views
FADRM: Fast and Accurate Data Residual Matching for Dataset Distillation

FADRM: Fast and Accurate Data Residual Matchin...

Computer Vision and Pattern Recognition
Avatar
librarian
365 views
HalluSegBench: Counterfactual Visual Reasoning for Segmentation
  Hallucination Evaluation

HalluSegBench: Counterfactual Visual Reasoning...

Computer Vision and Pattern Recognition
Avatar
librarian
449 views
Whole-Body Conditioned Egocentric Video Prediction

Whole-Body Conditioned Egocentric Video Prediction

Computer Vision and Pattern Recognition
Avatar
librarian
418 views
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven
  Thinking and Visual Drawing

Reinforcing Spatial Reasoning in Vision-Langua...

Computer Vision and Pattern Recognition
Avatar
librarian
527 views
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over
  Videos

Outside Knowledge Conversational Video (OKCV) ...

Computer Vision and Pattern Recognition
Avatar
librarian
407 views
Decoupling the Image Perception and Multimodal Reasoning for Reasoning
  Segmentation with Digital Twin Representations

Decoupling the Image Perception and Multimodal...

Computer Vision and Pattern Recognition
Avatar
librarian
557 views
Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via
  Spatial Reasoning

Direct Numerical Layout Generation for 3D Indo...

Computer Vision and Pattern Recognition
Avatar
librarian
572 views
Refer to Anything with Vision-Language Prompts

Refer to Anything with Vision-Language Prompts

Computer Vision and Pattern Recognition
Avatar
Shengcao Cao
529 views
Thinking with Generated Images

Thinking with Generated Images

Computer Vision and Pattern Recognition
Avatar
librarian
524 views
Let Androids Dream of Electric Sheep: A Human-like Image Implication
  Understanding and Reasoning Framework

Let Androids Dream of Electric Sheep: A Human-...

Computer Vision and Pattern Recognition
Avatar
Anastasia Kokkanen
566 views
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Delving into RL for Image Generation with CoT:...

Computer Vision and Pattern Recognition
Avatar
librarian
512 views
Let Androids Dream of Electric Sheep: A Human-like Image Implication
  Understanding and Reasoning Framework

Let Androids Dream of Electric Sheep: A Human-...

Computer Vision and Pattern Recognition
Avatar
librarian
541 views
SpatialScore: Towards Unified Evaluation for Multimodal Spatial
  Understanding

SpatialScore: Towards Unified Evaluation for M...

Computer Vision and Pattern Recognition
Avatar
Haoning Wu
522 views
VTBench: Evaluating Visual Tokenizers for Autoregressive Image
  Generation

VTBench: Evaluating Visual Tokenizers for Auto...

Computer Vision and Pattern Recognition
Avatar
librarian
579 views
Does Feasibility Matter? Understanding the Impact of Feasibility on
  Synthetic Training Data

Does Feasibility Matter? Understanding the Imp...

Computer Vision and Pattern Recognition
Avatar
librarian
508 views
MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal
  Mathematical Reasoning

MathCoder-VL: Bridging Vision and Code for Enh...

Computer Vision and Pattern Recognition
Avatar
librarian
586 views
StreamBridge: Turning Your Offline Video Large Language Model into a
  Proactive Streaming Assistant

StreamBridge: Turning Your Offline Video Large...

Computer Vision and Pattern Recognition
Avatar
librarian
542 views
Flow-GRPO: Training Flow Matching Models via Online RL

Flow-GRPO: Training Flow Matching Models via O...

Computer Vision and Pattern Recognition
Avatar
Jie Liu
757 views
DEIM: DETR with Improved Matching for Fast Convergence

DEIM: DETR with Improved Matching for Fast Con...

Computer Vision and Pattern Recognition
Avatar
huang shihua
607 views
DEIM: DETR with Improved Matching for Fast Convergence

DEIM: DETR with Improved Matching for Fast Con...

Computer Vision and Pattern Recognition
Avatar
huang shihua
556 views
HelloMeme: Integrating Spatial Knitting Attentions to Embed High-Level
  and Fidelity-Rich Conditions in Diffusion Models

HelloMeme: Integrating Spatial Knitting Attent...

Computer Vision and Pattern Recognition
Avatar
Songkey Z
693 views
Chat-Edit-3D: Interactive 3D Scene Editing via Text Prompts

Chat-Edit-3D: Interactive 3D Scene Editing via...

Computer Vision and Pattern Recognition
Avatar
shuangkang fang
702 views
Kvasir-VQA: A Text-Image Pair GI Tract Dataset

Kvasir-VQA: A Text-Image Pair GI Tract Dataset

Computer Vision and Pattern Recognition
Avatar
Sushant Gautam
708 views
3D modelling of survey scene from images enhanced with a multi-exposure
  fusion

3D modelling of survey scene from images enhan...

Computer Vision and Pattern Recognition
Avatar
DIEGO FRANCISCO GARCIA MOLINA
753 views
High-level camera-LiDAR fusion for 3D object detection with machine
  learning

High-level camera-LiDAR fusion for 3D object d...

Computer Vision and Pattern Recognition
Avatar
DIEGO FRANCISCO GARCIA MOLINA
690 views