Joonghyuk Shin

You can also call me Alex / 신중혁

joonghyuk AT snu.ac.kr

Joonghyuk Shin

Hi! I am a CS PhD student at SNU. I’m interested in building fast, interactive, and stateful generative world models for precise simulation and content creation. Recently, I have been working with Xun Huang on how AR video generative models can learn and preserve world state. Feel free to reach out for discussions or collaborations!

News

2026/06
Started my internship at NVIDIA in Santa Clara. Also excited to share that two papers I spent a lot of effort on, Text2TactileGraphics from my CMU visit and JAM-Flow from earlier last year, were accepted to ECCV 2026. Huge thanks to my co-authors!
2026/01
Two papers (MotionStream and DRPose) are accepted to ICLR 2026! + MotionStream is accepted as an Oral!
2025/12
I will be working at a stealth startup from Feb 2026 and at NVIDIA Spatial Intelligence Lab from Summer 2026, both on improving video world models.
2025/11
MotionStream is on arXiv. Gave talks at Pika AI, Daydream AI, and a stealth startup.
2025/06
A paper about text-based image editing is accepted in ICCV 2025.
2024/10
I will be joining Adobe Research in San Francisco as an intern on Eli Shechtman's team in Summer 2025, followed by a research visit to CMU's Robotics Institute in Fall 2025, hosted by Jun-Yan Zhu.

Education

Experience

Publications

* Equal contribution, † Equal advising

Can Video World Models Track Unobserved World States?
Joonghyuk Shin, Yicong Hong, Jaesik Park, Xun Huang
arXiv 2026
[Paper | Project Page | Code | ] We test whether video world models actually maintain hidden world state using an action-conditioned video Shell Game: standard backbones render plausible frames yet fall to chance beyond the training horizon, while negative-eigenvalue linear attention and nonlinear TTT fast weights extrapolate by carrying and revising a state across chunks.

Text-based Tactile Graphics Generation for the Visually Impaired
Ruihan Gao*, Joonghyuk Shin*, Ava Pun, Jaesik Park, Wenzhen Yuan, Jun-Yan Zhu
ECCV 2026
[Paper | Project Page | Code | Models | Data | ] We introduce Text2TactileGraphics, a fabrication-aware generative system that turns natural-language prompts into 3D-printable 2.5D tactile graphics with global geometry, tactile textures, and braille for blind and low-vision users.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
Mingi Kwon*, Joonghyuk Shin*, Jaeseok Jeong, Jaesik Park†, Youngjung Uh
ECCV 2026
[Paper | Project Page | ] We present a unified framework that jointly generates synchronized facial motion and speech using flow matching and MM-DiT, enabling diverse audio-visual synthesis tasks within a single model.

Direct Reward Fine-Tuning on Poses for Single Image to 3D Human in the Wild
Seunguk Do, Minwoo Huh, Joonghyuk Shin, Jaesik Park
ICLR 2026
[Paper | Project Page | Code | ] We introduce DRPOSE, a direct reward fine-tuning method that improves pose accuracy in single-view 3D human reconstruction, especially for dynamic and acrobatic poses, without requiring expensive 3D human assets.

MotionStream: Real-Time Video Generation with Interactive Motion Controls
Joonghyuk Shin, Zhengqi Li, Richard Zhang, Jun-Yan Zhu, Jaesik Park, Eli Shechtman, Xun Huang
ICLR 2026 (🔥 Oral, 1.13%)
[Paper | Project Page | Post | Code | ] MotionStream is a streaming (causal, real-time, and long-duration) video generation system with motion controls, operating at ~30 FPS on a single H100 GPU, unlocking new possibilities for interactive content generation.

Exploring MM-DiT

Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
Joonghyuk Shin, Alchan Hwang, Yujin Kim, Daneul Kim, Jaesik Park
ICCV 2025
[Paper | Project Page | Code | ] We perform a systematic analysis of MM-DiT's bidirectional attention mechanism and introduce a robust prompt-based editing method working across diverse MM-DiT models (SD3 series and Flux).

InstantDrag: Improving Interactivity in Drag-based Image Editing
Joonghyuk Shin, Daehyeon Choi, Jaesik Park
SIGGRAPH Asia 2024
[Paper | Project Page | Code (230+) | ] We present InstantDrag, an optimization-free pipeline for fast, interactive drag-based image editing that requires only an image and drag instruction as input, learning from real-world video datasets.

Fill-Up

Fill-Up: Balancing Long-Tailed Data with Generative Models
Joonghyuk Shin, Minguk Kang, Jaesik Park
arXiv 2023
[Paper | Project Page | ] We propose a two-stage method for long-tailed (LT) recognition using textual-inverted tokens to synthesize images, achieving SOTA results on standard benchmarks when trained from scratch.

StudioGAN

StudioGAN: A Taxonomy and Benchmark of GANs for Image Synthesis
Minguk Kang, Joonghyuk Shin, Jaesik Park
TPAMI 2023
[Paper | Code (3500+) | ] We present StudioGAN, a comprehensive library for GANs that reproduces over 30 popular models, providing extensive benchmarks and a fair evaluation protocol for image synthesis tasks.

Personal

Last updated on June, 2026 · with Face Looker