IT之家 · 智能时代 · 10/7/2026, 11:36:04 PM
Ex-ByteDance Intern Raises $30M for World Model Startup at $200M Valuation
Tian Keyu, former ByteDance intern and lead author of a NeurIPS 2024 best paper, has founded a world model startup valued at $200 million after raising nearly $30 million from investors including 5Y Capital. The team is developing a symbolic language system for visual processing, claiming to reduce video generation costs by 90%, with a full model release planned for 2027.
SOURCE COVERAGEOriginal coverage
On October 8, Bloomberg reported in the morning that Tian Keyu, a former ByteDance intern, has entered the world model startup space. His company has not yet made its public debut—lacking even a name or product—but has secured backing from prominent venture capital firms such as Five Yon Capital, positioning itself to compete with AI pioneers like Fei-Fei Li in this emerging field.
Tian graduated with a master’s degree from Peking University last summer (IT Home note: As stated in the original Bloomberg report; Tian is currently a doctoral student at Peking University). His team has raised nearly 200 million post-money** (approximately RMB 1.342 billion at current exchange rates).

▲ Tian Keyu's GitHub profile
Over the past year, he has focused his research on world models. AI scholars such as Fei-Fei Li and Yann LeCun are advancing this technical approach, aiming to enable AI to simulate the physical laws and operational mechanisms of the real world, thereby applying it to fields such as robotics, interactive video, and autonomous driving.
The ability to secure substantial investment before naming the company or launching a product reflects a new trend in AI investing: capital is increasingly seeking out and betting early on promising R&D talent. Last month, AMD acquired World Labs, founded by Fei-Fei Li, for billions of dollars. AMD CEO Lisa Su described the deal as a move to "acquire world-class talent."
Tian was the first author of a Best Paper award winner at NeurIPS 2024. The paper proposed a more efficient image generation method, allowing models to generate images step-by-step by predicting the next set of pixel blocks, similar to how large language models predict subsequent content.
Tian aims to apply the technology from his award-winning paper to his startup, designing a symbolic language specifically for processing visual content for AI. He believes that the information contained in videos is difficult to express efficiently using human language. The startup currently has a team of 10, including other former ByteDance employees.
According to Tian, the team has developed a system resembling an AI "dictionary," containing 200,000 symbols unreadable by ordinary humans, to help models process video. "Human language can never describe a video segment with complete accuracy. We are first creating a language for AI, then feeding it 100 million hours of video."
He claims this approach can reduce the cost of generating one second of video by at least a factor of ten. Before officially releasing the full model in 2027, he plans to demonstrate the technology's capabilities through live events. At this stage, the team is prioritizing technical development rather than focusing on consumer products or business models.