InTowards AIbyMengliu Zhao·4d agoPaper Walkthrough — Geometrically-Constrained Agent for Spatial ReasoningHow a formal task constraint bridges the semantic-to-geometric gap in spatial reasoning VLMs
InTowards AIbyMengliu Zhao·Jul 4Paper Walkthrough — MACT: A Multi-Agent Collaboration Framework for Visual Document UnderstandingFrom one model doing everything to four specialists doing one thing well
InTowards AIbyMengliu Zhao·Jun 29Paper Walkthrough — U-Mind: A Unified Framework for Real-Time Multimodal Interaction with…Can a single model think, talk, gesture, and render video simultaneously, while knowing how to reason?A response icon1A response icon1
Mengliu Zhao·Jun 21Paper Walkthrough — ReFAct: Empowering Multimodal Web Agents with Visual and Context FocusingFrom passive screenshot consumers to active visual reasoners
InTowards AIbyMengliu Zhao·Sep 25, 2025Explaining Tongyi DeepResearchToward the Era of Synthesized Data Training
Mengliu Zhao·Mar 31, 2025ML Engineering 201: Describing “Measurable Impact” from an IC’s PerspectiveUse the SWOP(S) method to measure your impact
InAI AdvancesbyMengliu Zhao·Mar 1, 2025Sparse TransformersFrom naive sparse attention to Kimi’s ultra-long context model and DeepSeek’s NSA
InAI AdvancesbyMengliu Zhao·Feb 5, 2025Inside DeepSeek V3A high-level overview of the technologies used in the DeepSeek v3 model
InTDS ArchivebyMengliu Zhao·Dec 24, 20242024 Survival Guide for Machine Learning Engineer InterviewsA year-end summary for junior-level MLE interview preparationA response icon3A response icon3