|
Course format
The 15 sessions consist of one introductory session, six lectures, and eight sessions of student presentation and discussion.
Introduction (1 session). Course overview and the policy for the presentation and discussion.
Lectures (6 sessions). Two lecture series of three sessions each.
(3 sessions) Multimodal Foundation Models (Sangdoo Yun): architectures and training of multimodal representations, multimodal foundation models, and a review of the recent world model literature.
(3 sessions) Multimodal Generation (Jin-Hwa Kim): foundations of 3D representation and rendering, 3D/4D generation, and a review of the recent physical AI literature.
Student presentations and discussion (8 sessions). Each session runs 3–4 sets, and a set consists of a 15–25 minute presentation followed by a 15-minute discussion. Three sets fill roughly 120 minutes and four sets roughly 160 minutes of the 170-minute slot, with a short break between sets.
|
|
Presentation
The student presentations are the core of this course, not a supplement to the lectures. Each student presents once, and the presentation is expected to reach the standard of a conference-style talk rather than a summary of the abstract.
Each student presents one paper, or a tightly coupled pair of papers, selected from the topic pool announced in Week 1. Topics are assigned by preference ranking; ties are broken on a first-come, first-served basis.
The presentation should cover: the problem and why it matters, the background needed to follow the method, the technical core of the method, the experimental evidence and what it does not establish, and the presenter's own critique.
Slides are submitted to the instructors no later than 2 days before the presentation date. A revised version, incorporating what came out of the discussion, is uploaded to the course repository within 3 days after the presentation.
|
|
Discussion
A 15-minute discussion collapses into silence, or into a conversation between one instructor and one student, unless it is given a structure. The following protocol is used for every set.
Panel. Two students are assigned in advance as the discussion panel for each presentation. Over the semester, each student accordingly gives one presentation and serves twice on a panel. The instructors moderate and take part in the discussion.
Preparation. Panelists may read the paper in advance, and are encouraged to do so when the topic is far from their own area. The discussion is nevertheless meant to build on what the presenter actually said.
The panel has two tasks.
Reviewer. Interrogate the work as a reviewer would — what the paper establishes and what it does not, weaknesses in the method, missing baselines or ablations, and claims that outrun the evidence.
Research-direction brainstorming. Work with the presenter, over the remainder of the 15 minutes, on where the line of work could go next.
Questions from the floor are welcome throughout, but the panel's questions take precedence. The moderator opens the floor whenever the panel's line of questioning is exhausted.
Panel report. Within one week of the session, each panelist submits a short report (1–2 pages) covering a summary of the presented work, the substance of the discussion, and the follow-up research directions that came out of it. Reports will be uploaded to the class webpage for the class, and any student is free to pursue any direction recorded in them, including for their own research.
Norms. Criticism is addressed to the paper and never to the presenter. Questions that begin from a misunderstanding are welcome; in a seminar on a moving research frontier, the misunderstandings are usually shared by half the room.
Example structure of the 15 minutes. The following is one way a set may run, offered as an illustration rather than as a rule.
0–3 min. Panelist A opens as reviewer: what the paper convincingly establishes, what it does not, and two prepared questions.
3–7 min. Panelist B follows up on the weaknesses; questions from the floor if the panel is done.
7–13 min. Panel and presenter brainstorm follow-up directions together; the instructors and the floor join in.
13–15 min. An instructor names the one or two most promising directions raised and states what a first experiment testing them would look like.
|
|
Grading
|
|
Attendance (10%)
Students who miss more than one-third of the classes without a valid excuse will receive 0 points. Otherwise, full credit will be given.
(Exceptions can be made when the cause of absence is deemed unavoidable by the course instructor.)
|
|
Attitude (10%)
Asking a thoughtful question during the lecture or participating in the discussion as a floor audience earns 1–2 points.
|
|
Presentation (50%)
Technical accuracy and depth, clarity of delivery and slides, and the quality of the presenter's own critique and open questions.
|
|
Discussion and participation (30%)
Service on two discussion panels and the resulting panel reports, together with questions and contributions from the floor in the other sessions.
* No midterm or final examination is planned for this course.
|
|
Textbooks
There is no official textbook. The reading list is the topic pool of recent conference and preprint papers announced in Week 1.
|
|
FAQ
Can I audit the course?
Please talk to the course instructors in person during the first week of class. Note that the enrollment cap is set by the number of presentation slots (24–32 students), since every enrolled student gives one presentation and serves on two panels.
The course seems to require proficiency in the Korean language to listen to talks, but my Korean isn't quite good. Maybe I shouldn't take it?
Just so you know, while the class content is primarily in English, the instruction will be in Korean. If this is challenging, please feel free to explore other class options.
I have no background in multimodal learning. Can I still take the course?
Prior knowledge of machine learning and deep learning is assumed. If multimodal learning itself is new to you, we encourage you to skim the Fall 2025 lecture slides on multimodal representation learning, multimodal foundation models, and neural graphics before the semester begins.
|
|