Skip to content

Jinho Choi

Ph.D. Candidate, KAIST AI

Visiting Researcher, Helmholtz Munich

Munich, Germany

Open to research internships and collaborations in the Munich area.

I am a Ph.D. candidate at KAIST and a visiting researcher at Helmholtz Munich, working toward safe and trustworthy AI through mechanistic interpretability and data-centric AI: making the concepts inside foundation models explicit and reproducible, and tracing model behavior back to the data it was trained on.

In my Ph.D. I use sparse dictionary learning to open up vision foundation models. With PatchSAE I trained sparse autoencoders on the CLIP vision transformer and found that adapting the model to a new task mostly remaps concepts it already has rather than learning new ones. ConceptScope turns the same lens on the data, breaking image datasets into interpretable concepts to expose biases no one had reported, and VisualScratchpad carries concept analysis into large vision–language models at inference time. Most recently I have been making the extracted concepts themselves reliable, so that analyses built on them hold up from one training run to the next.

I have also spent five years as an AI scientist at two startups, taking models from research prototype to deployed product. At Tomocube (2019–2022) I built 3D segmentation models for label-free holotomography of live cells that shipped in the company's commercial analysis software, plus a human-in-the-loop annotation tool that became its in-house labeling system. At Genesislab (2022–2024) I built the fairness pipeline for an AI video-interview product used by 150+ companies, cutting gender bias by 35% with no loss in accuracy; developed the talking-face avatar that serves as its AI interviewer; and built LLM conversational agents for a creator-persona app with 500K+ downloads.

During my M.S. at Korea University I worked where machine learning meets human–computer interaction: topic models and visual analytics that detect and explain local events in social media streams (STExNMF, TopicOnTiles), and making charts accessible to people with visual impairments (Visualizing for the Non-Visual).

Recent

  • Apr 2026 VisualScratchpad presented at the ICLR 2026 Workshop on Trustworthy AI.
  • Feb 2026 Started as a visiting researcher at Helmholtz Munich (Dynamical Inference Lab, PI: Steffen Schneider).
  • Sep 2025 ConceptScope accepted to NeurIPS 2025.
  • Jan 2025 PatchSAE accepted to ICLR 2025.

Research

Mechanistic interpretability

Sparse dictionary learning for reproducible concept discovery in foundation models.

Safety & trustworthy AI

Fairness diagnosis and mitigation, and concept-level auditing of model behavior.

Data-centric AI

Characterizing dataset bias through model internals, and human-in-the-loop annotation.

Publications

marks work I would point to first.

  1. ConSDL
    2026

    Beyond SAEs: Consistent Concept Discovery with Sparse Coding

    Jinho Choi, Hyesu Lim, Jaegul Choo, Steffen Schneider

    Replaces sparse autoencoders with sparse coding so the concepts extracted from a model come out the same, run after run.

    Under review

  2. VisualScratchpad
    2026

    VisualScratchpad: Grounding Visual Concepts in Large Vision Language Models

    Hyesu Lim, Jinho Choi, Taekyung Kim, Byeongho Heo, Jaegul Choo, Dongyoon Han

    Grounds visual concepts inside large vision–language models at inference time, showing what the model looks at when it answers.

    ICLR 2026 Workshop on Trustworthy AI paper

  3. Teaser figure for ConceptScope 2025

    ConceptScope: Characterizing Dataset Bias via Disentangled Visual Concepts

    Jinho Choi, Hyesu Lim, Steffen Schneider, Jaegul Choo

    Breaks image datasets into interpretable visual concepts and measures how each spreads across classes, surfacing previously unreported biases.

    NeurIPS 2025 paper code project page

  4. Teaser figure for PatchSAE 2025

    Sparse Autoencoders Reveal Selective Remapping of Visual Concepts During Adaptation

    Hyesu Lim, Jinho Choi, Jaegul Choo, Steffen Schneider

    A patch-level sparse autoencoder on CLIP shows that adaptation mostly remaps existing visual concepts rather than learning new ones.

    ICLR 2025 paper code project page

  5. Slice & Conquer
    2024

    Slice and Conquer: A Planar-to-3D Framework for Efficient Interactive Segmentation of Volumetric Images

    Wonwoo Cho, Dongmin Choi, Hyesu Lim, Jinho Choi, Saemee Choi, Hyun-seok Min, Sungbin Lim, Jaegul Choo

    Lifts a few labeled 2D slices into full 3D masks and asks for correction only where the model is uncertain, speeding up volumetric annotation.

    WACV 2024

  6. Fair AVI
    2023

    Fairness-aware Multimodal Learning in Automatic Video Interview Assessment

    Changwoo Kim, Jinho Choi, Jongyeon Yoon, Daehun Yoo, Woojin Lee

    Diagnoses and mitigates demographic bias in multimodal video-interview scoring without giving up accuracy.

    IEEE Access 2023

  7. Label-free 3D
    2021

    Label-free Three-dimensional Analyses of Live Cells with Deep-learning-based Segmentation Exploiting Refractive Index Distributions

    Jinho Choi, Hye-Jin Kim, et al.

    Segments organelles and whole cells in 3D from refractive-index tomograms of live cells, with no fluorescent labeling.

    bioRxiv 2021

  8. 3D Cell Inst.
    2021

    3D Cell Instance Segmentation via Point Proposals using Cellular Components

    Jinho Choi, Junwoo Park, Hyun-seok Min, Hyungjoo Cho, Sungbin Lim, Jaegul Choo

    Separates touching cells in 3D microscopy volumes by proposing cell points from their cellular components.

    SPIE 2021

  9. Non-Visual Vis
    2019

    Visualizing for the Non-Visual: Enabling the Visually Impaired to Use Visualization

    Jinho Choi, Sanghun Jung, Deokgun Park, Jaegul Choo, Niklas Elmqvist

    Extracts the underlying data from chart images so screen-reader users can explore visualizations.

    Computer Graphics Forum (EuroVis) 2019

  10. TopicOnTiles
    2018

    TopicOnTiles: Tile-based Spatio-Temporal Event Analytics via Exclusive Topic Modeling on Social Media

    Minsuk Choi, Sungbok Shin, Jinho Choi, et al.

    Tile-based visual analytics that surfaces local events in social media through spatio-temporally exclusive topics.

    CHI 2018

  11. STExNMF
    2017

    STExNMF: Spatio-Temporally Exclusive Topic Discovery for Anomalous Event Detection

    Sungbok Shin, Minsuk Choi, Jinho Choi, et al.

    A non-negative matrix factorization that finds topics exclusive in space and time to flag anomalous events in social media.

    IEEE ICDM 2017

Experience

Feb 2026 — Present

Munich, Germany

Helmholtz Munich

Visiting Researcher · Dynamical Inference Lab (PI: Steffen Schneider), joint project with KAIST DAVIAN Lab

  • Developed ConSDL, which extracts the concepts a model has learned reproducibly. With sparse autoencoders, the standard tool, two training runs can disagree on what the concepts are, undermining any analysis built on them. ConSDL recovers the same concepts run after run and is the most consistent method on vision and language models ConSDL.

Aug 2022 — Present

Daejeon, Korea

KAIST, DAVIAN Lab

Ph.D. Researcher

  • Developed ConceptScope, which breaks a dataset into human-interpretable visual concepts and measures how each concept distributes across classes, uncovering many previously unreported biases in real-world image datasets ConceptScope.
  • Trained sparse autoencoders on the CLIP vision transformer (PatchSAE), discovering 49K localized visual concepts that explain how the model reaches its predictions and how that changes under adaptation to new tasks PatchSAE. Extended the analysis to large vision–language models, grounding visual concepts at inference time (VisualScratchpad) VisualScratchpad.

Sep 2022 — Nov 2024

Seoul, Korea

Genesislab

AI Scientist

  • Built the fairness pipeline for viewinterHR, a video interview assessment product used by 150+ companies: a diagnosis stage that quantifies how sensitive and nuisance attributes affect outcomes, and a mitigation algorithm that reduces group disparity in features and outputs, cutting gender bias by 35% with no loss in accuracy Fair AVI.
  • Developed a talking-face generation method with accurate lip-sync and high video clarity, deployed as the AI interviewer avatar in viewinterHR. Worked with HR experts to ensure the generated interviewer reflected real interview dynamics.
  • Built the conversational agents behind Zuicy, LLM personas of YouTube creators (500K+ downloads), using prompt chaining, retrieval-augmented generation, and memory recall, plus an automated pipeline for building each creator persona.

Feb 2019 — Sep 2022

Seoul, Korea

Tomocube

AI Researcher

  • Developed 3D cell segmentation models for four organelle types and cell instances from label-free refractive index images, achieving state-of-the-art accuracy, and integrated them into commercial cell analysis software 3D Cell Inst. Label-free 3D.
  • Built a human-in-the-loop segmentation model that lifts labeled 2D slices to 3D masks and actively requests correction on uncertain slices, improving annotation speed and accuracy by 9.5% and serving as the in-house annotation system Slice & Conquer.

Education

  • 2022 — 2028

    Ph.D. in Artificial Intelligence

    KAIST · Advisor: Jaegul Choo · expected Feb 2028 · part-time 2022–2024

  • 2017 — 2019

    M.S. in Computer Science and Engineering

    Korea University · Advisor: Jaegul Choo

  • 2013 — 2017

    B.S. in Computer Science and Engineering

    Korea University

Service

  • Reviewer · NeurIPS, ICLR