# Zhang Xiangyu's Talk on Multimodality & Reasoning

A truly inspirational view on CV -> Multimodality -> Reasoning

Canonical: https://astro-blog-bice-chi.vercel.app/blog/image-reasoning

Edition: en · Language: en · Revision: 2026-09-21

Published: 2025-07-05

Updated: 2026-09-21

Sources checked and uncertain recollections identified; the argument remains a personal reading of the July 2025 talk.

Updated   2026-09-21

This is an interview with Zhang Xiangyu, a long-time research scientist. This is not a traditional peer-reviewed scientific paper per se; in my opinion, it’s a more elaborate, Ilya-level talk[^1]. I highly recommend listening/reading it in a very intellectual mode/environment—it’s worth your time.

> Podwise[^2] already has many highlights and summaries automatically generated. So I am not going to copy and paste here. However, I will list the most important lessons I drew from the talk and my fuzzy thoughts relevant to it.

## [1\. Image alone scales poorly](https://astro-blog-bice-chi.vercel.app/blog/image-reasoning#1-image-alone-scales-poorly)

This is actually a bigger claim than the original talk, which focused more on “self-supervised learning for images scales poorly”.

First of all, there is **NO SUCH THING** as **UNSUPERVISED** learning. LLMs have strong supervision in two forms: 1) language itself is well-structured, and 2) selecting which data to train on is a form of annotation.[^3]

Zhang’s insight is that SSL in vision is actually supervised and hence limited by the developers’ knowledge of crafting augmentations. More specifically:

*   **Contrastive Learning** is essentially learning a **HUMAN’S** understanding of object invariance regarding color, scale, aspect ratio, etc.
*   A **Masked Autoencoder** is essentially learning occlusion invariance

But there’s more we would hope for: physics, composition, object relationships, causal effects I know I'm throwing some non-orthogonal concepts in here; I'm just lacking better words

SSL in video may have a shot at solving this, but temporal redundancy makes brute-force scaling look inefficient to me. VideoMAE uses very high masking ratios to exploit that redundancy[^4]; this does not establish a general limit on video learning. More efficient solutions are yet to come.

Now comes my personal take on the “supervised learning **ALSO** scales poorly” statement. Though Google DeepMind[^5], Meta[^6], and BAAI[^7] also show improvements with **MASSIVE** scaling on vision encoders, these improvements are, in my opinion, somewhat nice-to-have. I doubt they bump many applications from the not-okay → okay → great tier of usage. I even believe image resolution (and hence the number of tokens utilized) plays a bigger role in image understanding - at least it results in significant differences in the holistic understanding vs localization/OCR trade-off.[^8]

## [2\. The LLM inverse scaling phenomenon](https://astro-blog-bice-chi.vercel.app/blog/image-reasoning#2-the-llm-inverse-scaling-phenomenon)

This is a concurrent insight from the back-and-forth [^9] discussion [^10] on “emergent abilities,“[^11] and Zhang’s insight is more intuitive to follow: when LLMs get larger, they **DO** have better intuition for solving (reasoning) tasks, hence they prefer to take the shortcut of spelling out answers directly, which incurs higher error rates than faithfully solving problems step-by-step. When properly instructed (CoT), larger models **ARE** better.

This makes me wonder if vision can be solved with data and compute at all. After all, the native vision model does not have **DYNAMIC** “reasoning circuits” that can reason with more compute via CoT. **NOTE**: I recall an example of reasoning through autoregressive image prediction, but have not recovered the intended source. I leave that possibility as a question, not evidence for this claim.

## [3\. The power of a reasoning model (with RL)](https://astro-blog-bice-chi.vercel.app/blog/image-reasoning#3-the-power-of-a-reasoning-model-with-rl)

This is the most fascinating part of the talk. It is no secret:

*   **WHAT**: RL can improve generalization in the studied tasks[^12]
*   **HOW**: DeepSeek-R1’s use of GRPO[^13] makes it widely popular
*   **BUT WHY?** It’s too simple to be true—why didn’t it happen earlier?

My reading of Zhang’s hypothesis is that much of the reasoning pattern is already present in pre-training, and post-training helps elicit it. I do not take this as proof that post-training can introduce **NO NEW KNOWLEDGE**. RL simply does a better job of eliciting **EXISTING PATTERNS** (like CoT) from pre-training, and the “aha moment” words (like “wait…”) are common phrases in high-quality math discussion data. When combined with test-time compute, which allows models to dynamically leverage more tokens on “thinking paths” with higher success rates, and also to adapt from previously wrong choices.

My intuition is that:

*   Instead of slamming hard on tokens that have a multimodal nature and non-reducible variance, we should supervise on thinking circuits
*   RL is a better way (than SFT) to search and elicit existing patterns
*   In layman’s terms, RL explores better solutions from the model’s point of view RL term 'on-policy' and helps the model develop a better “intuition” in the solution space instead of the token space again, this is my intuition, and guess how I developed my intuition?🤔

> Update: I found this fascinating explanation of WHY RL works[^14]

![why RL is promising?](https://astro-blog-bice-chi.vercel.app/assets/blog/image-reasoning/RL_is_promising.webp)

## [4\. The path to reasoning with images](https://astro-blog-bice-chi.vercel.app/blog/image-reasoning#4-the-path-to-reasoning-with-images)

OpenAI’s O3 is a living testament[^15] that reasoning with images is possible. If you haven’t seen it, be sure to watch it—truly special, even magical.

Just to be clear, O3’s image implementation is still largely **UNKNOWN**.

Zhang’s insight is that it **PROBABLY** can be traced back to pre-training as well:

*   There are lots of explanation patterns for images
*   These patterns often involve annotation or simply zooming to regions of interest, e.g., fixing electronics, labeling the special usage of tools, etc.
*   O3’s thinking patterns can be seen as successful elicitation of cropping/zooming/annotation patterns
*   As a side note, image manipulation is implemented with live coding in Python at least it seems that way ; though I think more structured, parameterized function calls would be more appropriate.

Sources & further reading

## References

[^1]: Why Next-Token Prediction Could Surpass Human Intelligence: [video](https://www.youtube.com/watch?v=Yf1o0TQzry8)

[^2]: [https://podwise.ai/dashboard/episodes/4209997](https://podwise.ai/dashboard/episodes/4209997)

[^3]: Unverified recollection: I have not identified the original talk or speaker, so this is not a verified attribution.

[^4]: VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training: [https://arxiv.org/abs/2203.12602](https://arxiv.org/abs/2203.12602)

[^5]: SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features: [https://arxiv.org/abs/2502.14786](https://arxiv.org/abs/2502.14786)

[^6]: Perception Encoder: The best visual embeddings are not at the output of the network: [https://arxiv.org/abs/2504.13181](https://arxiv.org/abs/2504.13181)

[^7]: EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters: [https://arxiv.org/abs/2402.04252](https://arxiv.org/abs/2402.04252)

[^8]: Related evidence on resolution and representation trade-offs: Perception Encoder, [https://arxiv.org/abs/2504.13181](https://arxiv.org/abs/2504.13181). This paper does not establish that resolution matters more than model size in general.

[^9]: Inverse Scaling: When Bigger Isn’t Better: [https://arxiv.org/abs/2306.09479](https://arxiv.org/abs/2306.09479)

[^10]: Inverse scaling can become U-shaped: [https://arxiv.org/pdf/2211.02011](https://arxiv.org/pdf/2211.02011)

[^11]: [https://www.jasonwei.net/blog/common-arguments-regarding-emergent-abilities](https://www.jasonwei.net/blog/common-arguments-regarding-emergent-abilities)

[^12]: SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training: [https://arxiv.org/abs/2501.17161](https://arxiv.org/abs/2501.17161). Its results concern the tested rule and visual variants, not a universal law about SFT.

[^13]: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning: [https://arxiv.org/abs/2501.12948](https://arxiv.org/abs/2501.12948)

[^14]: \[UCLA RL-LLM\] Chapter 0: Course outline and prologue: [https://www.youtube.com/watch?v=q9972BRoXzQ](https://www.youtube.com/watch?v=q9972BRoXzQ)

[^15]: [https://simonwillison.net/2025/Apr/26/o3-photo-locations/](https://simonwillison.net/2025/Apr/26/o3-photo-locations/)
