
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
About this episode
🤗 Upvotes: 90 | cs.LG, cs.CL, cs.CV
Authors:
Yicheng Zou, Dongsheng Zhu, Lin Zhu, Tong Zhu, Yunhua Zhou, Peiheng Zhou, Xinyu Zhou, Dongzhan Zhou, Zhiwang Zhou, Yuhao Zhou, Bowen Zhou, Zhanping Zhong, Zhijie Zhong, Haiteng Zhao, Penghao Zhao, Xiaomeng Zhao, Zhiyuan Zhao, Yechen Zhang, Jin Zhang, Wenwei Zhang, Hongjie Zhang, Zhuo Zhang, Wenlong Zhang, Bo Zhang, Chao Zhang, Chen Zhang, Yuhang Zang, Fei Yuan, Jiakang Yuan, Jiashuo Yu, Jinhui Yin, Haochen Ye, Qian Yao, Bowen Yang, Danni Yang, Kaichen Yang, Ziang Yan, Jun Xu, Yicheng Xu, Wanghan Xu, Xuenan Xu, Chao Xu, Ruiliang Xu, Shuhao Xing, Long Xing, Xinchen Xie, Ling-I Wu, Zijian Wu, Zhenyu Wu, Lijun Wu, Yue Wu, Jianyu Wu, Wen Wu, Fan Wu, Xilin Wei, Qi Wei, Bingli Wang, Rui Wang, Ziyi Wang, Zun Wang, Yi Wang, Haomin Wang, Yizhou Wang, Lintao Wang, Yiheng Wang, Longjiang Wang, Bin Wang, Jian Tong, Zhongbo Tian, Huanze Tang, Chen Tang, Shixiang Tang, Yu Sun, Qiushi Sun, Xuerui Su, Qisheng Su, Chenlin Su, Demin Song, Jin Shi, Fukai Shang, Yuchen Ren, Pengli Ren, Xiaoye Qu, Yuan Qu, Jiantao Qiu, Yu Qiao, Runyu Peng, Tianshuo Peng, Jiahui Peng, Qizhi Pei, Zhuoshi Pan, Linke Ouyang, Wenchang Ning, Yichuan Ma, Zerun Ma, Ningsheng Ma, Runyuan Ma, Chengqi Lyu, Haijun Lv, Han Lv, Lindong Lu, Kuikun Liu, Jiangning Liu, Yuhong Liu, Kai Liu, Hongwei Liu, Zhoumianze Liu, Mengjie Liu, Ziyu Liu, Wenran Liu, Yang Liu, Liwei Liu, Kaiwen Liu, Junyao Lin, Junming Lin, Tianyang Lin, Dahua Lin, Jianze Liang, Linyang Li, Peiji Li, Zonglin Li, Zehao Li, Pengze Li, Guoyan Li, Lingkai Kong, Linglin Jing, Zhenjiang Jin, Feifei Jiang, Qian Jiang, Junhao Huang, Zixian Huang, Haian Huang, Zhouqi Hua, Han Hu, Linfeng Hou, Yinan He, Conghui He, Tianyao He, Xu Guo, Qipeng Guo, Aijia Guo, Yuzhe Gu, Lixin Gu, Jingyang Gong, Qiming Ge, Jiaye Ge, Songyang Gao, Jianfei Gao, Xinyu Fang, Caihua fan, Yue Fan, Yanhui Duan, Zichen Ding, Shengyuan Ding, Xuanlang Dai, Erfei Cui, Ganqu Cui, Pei Chu, Tao Chu, Guangran Cheng, Yu Cheng, Kai Chen, Yongkang Chen, Chiyu Chen, Guanzhou Chen, Qiaosheng Chen, Sitao Chen, Xin Chen, Haojiong Chen, Yicheng Chen, Weihan Cao, Yuhang Cao, Qinglong Cao, Lei Bai
Title:
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
Arxiv:
http://arxiv.org/abs/2603.25040v1
Abstract:
We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond stronger reasoning and image-text understanding capabilities, its intelligence is augmented with advanced agent capabilities. Simultaneously, its scientific expertise has been vastly expanded to master over 100 specialized tasks across critical science fields, including chemistry, materials, life sciences, and earth sciences. Achieving this massive scale is made possible by the robust infrastructure support of XTuner and LMDeploy, which facilitates highly efficient Reinforcement Learning (RL) training at the 1-trillion parameter level while ensuring strict precision consistency between training and inference. By seamlessly integrating these advancements, Intern-S1-Pro further fortifies the fusion of general and specialized intelligence, working as a Specializable Generalist, demonstrating its position in the top tier of open-source models for general capabilities, while outperforming proprietary models in the depth of specialized scientific tasks.
Get every episode summarized
Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
191 searchable segments. Every word is indexed and playable.
Full transcript
Daily Paper Cast — Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Hey, everyone. Welcome back to another exciting episode of the Daily Papercast. I'm Echo. And I'm Nova. Today we've got an intriguing paper lined up for discussion. It's going to be a deep dive into the world of large language models and their implications in science. We'll break it down into four segments, introduction, methods, experiments, and related work. Each part will take about four to five minutes. So, let's get started. All right. Let's dive right in. The title of today's paper is Pause for Effect. Intern S1 Pro, Scientific Multimodal Foundation Model at Trillion Scale. Quite a mouthful, huh? Absolutely. This paper is authored by a team from the Shanghai AI Laboratory, led by Jian Wen Wang and Jian Xin Lai, with Ji Kyeong Zhang as the last author. It's a monumental effort indeed. Let's kick things off with the introduction section. The paper starts by discussing the advent of large language models, LLMs, and visual language models, VLMs.
These models have revolutionized AI, offering unprecedented capabilities in reasoning, generation, and multimodal understanding. They've become critical tools in AI for science, AI for S, enabling breakthroughs in fields like protein structure prediction and materials design. Right. The idea is that these models serve as unified interfaces that process vast amounts of scientific literature, experimental data, and domain-specific knowledge. Essentially, they bridge gaps between different scientific disciplines, which is pretty mind-blowing when you think about it. Totally. The authors point out that the diversity in scientific domains necessitates scaling up the model size. For example, just like multi-lingual machine translation models require exponentially more parameters to handle multiple language bears. Scientific models need to handle the complexity and specificity of various scientific fields, like chemistry, biology, physics, and Earth sciences. And that's where intern S1 Pro comes into play.
This is the first one trillion parameter scientific multimodal foundation model. It scales to an unprecedented size, delivering comprehensive enhancements across general and scientific domains. Beyond stronger reasoning and image-text understanding, it has advanced agent capabilities, enabling it to autonomously plan and execute complex scientific workflows. Yes. It's designed to master over 100 specialized tasks across various critical science fields, including chemistry, materials, life sciences, and Earth sciences. The model's architecture allows it to outperform specialized models in several scientific tasks, by leveraging joint training on both general and specific tasks. This is counter to the common belief that specialized models are superior for niche tasks. The findings reveal that a sufficiently large generalist model, when trained jointly, can exceed the performance of specialized models. Exactly. The paper also mentions the sage framework, which stands for synergistic architecture for generalizable experts.
It incorporates three layers, foundation, fusion, and evolution. This framework is key to enabling synergistic improvements across domains. That's crucial because scaling up model parameters isn't just about adding more computational power. It introduces new challenges like training and stability, and the need for efficient optimization of router embeddings. The paper proposes solutions for maintaining both stability and efficiency, through mechanisms like group routing and expert parallelism. Right. And the engineering effort behind maintaining high training throughput is notable. By co-designing algorithms and infrastructure, they manage to scale the training framework, to support trillion scale models, while maintaining impressive efficiency. That's no small feat. Absolutely. The integration of general and specialized intelligence in intern S1 Pro not only fortifies its capabilities, but also sets a standard in the open source community. Outperforming some proprietary models in the depth of specialized scientific tasks.
So, to sum up the introduction section, intern S1 Pro represents a significant advancement in AI models, especially in the realm of scientific discovery. It's an impressive blend of general and specialized capabilities, making it a versatile tool for a wide range of scientific applications. We're just getting started, and it's already so fascinating. Stay tuned as we move on to the methods section next, where we'll dive into the technical details behind this groundbreaking model. You don't want to miss it. All right, Echo. Let's dive into the methods section of this paper. Absolutely, Nova. So, where do we begin? There's a lot to unpack. How about we start with the architecture? The paper elaborates on intern S1 Pro through an expert expansion process, incorporating something called grouped routing. Grouped routing? That sounds interesting. What exactly does it entail? Great question. In layman's terms, grouped rooting ensures that the computation load is balanced across different devices during training,
which is especially crucial when dealing with ultra-large scale MOE, that is, mixture of experts, models like intern S1 Pro. Ah, I see. Traditional routing strategies could lead to load imbalances, potentially causing spikes in memory usage. That doesn't sound very efficient. Exactly. To address this, they divided all experts into distinct groups based on device mapping. Only the top experts within each group are selected, which helps balance the load and stabilizes the training process significantly. Makes perfect sense. So how do they handle the data for training this beast of a model? They used 6 trillion tokens of high-quality multi-modal data. A significant portion of this was image text and pure text data. They even designed a specialized caption pipeline for scientific imagery, ensuring the quality and alignment of text with images. Imagine the kind of data curation efforts that must have gone into this.
Wow. That's extensive. Speaking of innovation, the paper also mentioned something called a straight-through estimator, or STE for sparse expert routing. What's that about? The STE is quite ingenious. Essentially, it preserves the standard sparse top-case selection during the forward pass. But during the backward pass, it allows gradients to flow through the full dense softmax distribution. This enables all experts to receive meaningful gradient updates, preventing what they call gradient sparsity, which can degrade the router's optimization capabilities. That's clever. So it essentially smooths out the training process, ensuring all experts get their fair share of updates. You got it. And to make it even more robust, they used a combined approach of FP8 quantization for reinforcement learning. This took care of memory pressure issues without sacrificing performance. FP8 quantization, huh? That's reducing the precision of some computations to only 8 bits, right? How does that help here?
Yes. SP8 quantization helps by significantly reducing the memory and computational footprint. They applied this sparsely, though, focusing on gamma operations primarily in expert layers, which are tolerable to reduced precision. More sensitive layers like the LM head kept higher precision to maintain fidelity and computations. Incredible. With all this optimization, I bet their data integration strategies for combining general and scientific data are worth mentioning, too. Definitely. They used three major strategies, structured scientific data transformation, scientific data diversification, and system prompt isolation. Tell me more about these strategies. How do they tackle the integration challenges? Sure. For structured scientific data transformation, they focused on converting highly structured tabular information into narrative text through template construction. This ensures semantic consistency and minimal information loss, aligning it nicely with general data representation.
Smart move. And how about scientific data diversification? This one is about preventing overfitting. They introduced prompt diversification by varying instructions for core scientific concepts. They also tackled overly simplistic outputs by generating complete reasoning chains through a rollout mechanism, turning simple knowledge recall into logical deduction. And system prompt isolation? Sounds like something serious. It is. This strategy involves injecting mutually exclusive system level prefixes during training for scientific and general data. It creates independent contextual environments, reducing conflicts and ensuring stability. Mixing different data types without such isolation would probably have made the model go, what on earth is happening? True. Laughs. Oh, and let's not forget the post training phase where they scaled reinforcement learning for this trillion parameter model. They introduced a mixed precision scheme to maintain stable training, despite the large model size.
Post training improvements must have been a task. How do they handle the sued memory requirements? They adopted FP8 quantization here too, focusing on reducing memory footprint where possible. But given the extreme sparsity, they had to manage the training inference consistency very carefully to avoid instability. All right. So all the advanced routing strategies, the multi-tier data processing, diverse data sets and a robust quantization approach, it's clear a lot went into making intern s1 pro as effective as it is. Absolutely. The methods outline not only show careful engineering, but also elegant solutions to complex problems. It's a brilliant blend of scalability, data finesse and computational optimization. And that wraps up our detailed look into the method section of this incredible paper. All right. Echo, let's dive into the meat of the paper, the experiments and results. This is where the real action happens, right?
Absolutely, Nova, what's a model without some solid experiments to back it up, eh? So the paper on intern s1 pro covers a wide range of benchmarks and comparisons. You ready to break it down? Definitely. Let's start with the evaluation setup they used. The team conducted extensive experiments using three main toolkits, open compass, VLM eval kit and agent compass. They used both thinking and non-thinking configurations to see how well intern s1 pro performs in different scenarios. Right. The thinking configuration involves more complex evaluation settings, which are based on the official settings for each benchmark. And the non-thinking configuration, on the other hand, seems more straightforward. It's designed to evaluate simpler tasks. They tested intern s1 pro on various scientific and general-purpose tasks, covering both text only and multimodal settings. It's like putting the model through a bootcamp with different challenges to see where it shines.
Exactly. They compared intern s1 pro with other state-of-the-art models and surprise, surprise. It performed exceptionally well. For example, in scientific reasoning tasks, intern s1 pro outperformed leading closed source models by a substantial margin. On the psi-reasonner benchmark, it scored 55.5 compared to Gemini 3 pros, 14.7 and GPT 5.2s, 13.6. Wow. That's quite a leap. And they didn't stop there. The model also achieved top positions on benchmarks like small-instruct, nap bench, mall instructions, biology instruction, and XLRS bench. These benchmarks test everything from molecular science to large-scale remote sensing imagery. It's like this model is a jack of all trades, but is it also a master of them? It seems like it, and for general-purpose tasks, intern s1 pro kept up its strong performance. For instance, on the Amy 2025 and MMLU Pro benchmarks, it scored 93.1 and 86.6 respectively.
It's like running a marathon with different terrains and still keeping the lead. Impressive indeed. Another cool aspect is the model's performance in the time series tasks. The psi-TS benchmark, which includes tasks like ASU01, BIU03, and EAU01, showed how the model's dedicated time series module really shines. For EAU01, intern s1 pro blew others out of the water with an F1 score of 99.5. Now, that's something. It seems like their adaptive sub-sampling process and dedicated time series encoder made a huge difference. Handling time series data effectively is no small feat. So, echo, how did intern s1 pro do in those specialized tasks like biological sequence and structure tasks? Oh, this is interesting. They did a case study comparing intern s1 pro to a specialized model biology instruction. The results were fascinating. For instance, in the protein fluorescence task, intern s1 pro scored a jaw-dropping 78.14 compared to biology instructions 2.57.
And on protein function EC, it scored 72.70 against 19.79. Talk about outperforming a specialized model. It's like having a Swiss Army knife that's also better than the single-purpose tools. Totally. These results suggest that integrating general and specialized capabilities can create a synergistic effect, making the model much more effective in professional domains. It's not just about having more tools in the box. It's about how they all work together. And what about training stability and efficiency? With such a large model, there must have been some hurdles, right? For sure. Scaling up to a trillion-parameter model brought new challenges. They tackled these by using techniques like FP8 quantization during the RL phase and a mixed precision training approach. They even introduced something called rollout router replay to ensure expert selection consistency between training and inference phases. That's quite innovative. Ensuring precision consistency between training and inference must have helped a lot in maintaining the model's robustness and performance.
Absolutely. Plus, they made use of a custom caption pipeline to generate high-quality image text pairs, which was crucial for the model's understanding of scientific visuals. It's like getting the best possible lens for your camera. Sharper images, better results. Sharpe as attack, echo. It seems like the team thought of everything. From diverse data input to robust training strategies, they covered all bases to make intern S1 Pro truly exceptional. Exactly, Nova. And that wraps up our deep dive into the experiments and results of this mind-blowing model. I hope our listeners are as pumped as we are about these innovations. I'm sure they are. That's the end of our experiment section, folks. Stay tuned for more exciting insights in the next segment. All right, Nova. Let's dive into the related work section. This is where things get interactive. You know what I mean? Absolutely, echo. In any research, seeing how it builds on existing work is just fascinating. It's like connecting the dots in a puzzle.
Right. So the first thing the paper touches on in this section is how previous models have grappled with the sheer complexity and diversity of scientific tasks. We're talking about tasks ranging from protein structure prediction to materials design. These areas require very specific domain knowledge and reasoning patterns that are vastly different from natural language processing alone. Ford or O. Wow, that's intense. And it makes sense that the paper would build on these foundational challenges. Instead of treating scientific data as just another language, it's addressing the unique languages of different scientific domains. Exactly. The authors point out that scientific data is often represented in structured formats, which are different from traditional text data. They discuss leveraging these structured data sets like pub chem for chemical data to improve the model's foundational understanding. That's pretty innovative by transforming these structured data into a format the model can understand. It's creating a bridge between raw scientific data and general language models.
Did they mention any specific previous work that has tackled this in a similar way? Yes, they reference several works. For example, they mention ice pop, which handles important sampling and masks tokens exhibiting large distribution shifts between training and inference. This kind of work is crucial for ensuring that the model remains stable and accurate during different phases of use. Consistency between training and inference is key. Otherwise, you're essentially training for a different game than you're playing. So does the paper also discuss how it improves on these previous methods? Yes, it does. One of the major improvements is the use of rollout router replay. This technique records which experts were used during the training rollout and replays those same experts during policy updates, ensuring that the model's behavior remains consistent from training to inference. That's pretty clever. It's like they're ensuring that the practice drills are as close to the actual game as possible. What about handling the complexity of multimodal data? How does this research build on previous work there?
Good question, Nova. They highlight some interesting previous approaches, including Fourier position encoding, FOPE, which is used to capture the continuous wave-like characteristics of physical signals. This is an evolution from traditional positional encoding methods like sinusoidal encodings, which were more geared towards text data. That's fascinating. Fourier position encoding sounds like a big leap. So is this approach novel or are they improving on an existing concept? It's an improvement on an existing concept. They reference works that have explored similar grounds, but point out that previous methods weren't adapted handling the continuous nature of signals in physical data. This paper combines the discrete particle nature of tokens with the continuous wave characteristics, bridging a critical gap. So it sounds like they're really stepping up the game in terms of how the model interprets different types of data. Did they discuss any previous benchmarks or models that tackled these issues but on a smaller scale?
Yes, definitely. They mentioned several benchmarks such as the MMLU Pro for a robust multi-task language understanding that demands deep reasoning across several subjects. They emphasize that their new model aims to improve on these benchmarks by integrating more comprehensive and scalable solutions. Makes sense. Comparing to benchmarks really gives a solid ground to evaluate progress. What about computational resources? Did they touch on any advancements in that area? Yes, indeed. They mention intricate techniques to handle the memory challenges that come with scaling up. For instance, they implemented a grouped router design to ensure balanced loads across devices, even with the massive parallel activations required by such models. Ah, the good old load balancing. It's the unsung hero of keeping large scale models efficient. What else is notable from their related work discussions? Probably the emphasis on mixed precision training approaches. They extend previously established methods like FPA quantization, which is pivotal for efficient reinforcement learning in their model.
You know, it's really impressive how they're integrating and improving upon such a wide array of existing technologies. It's like they're crafting the ultimate Swiss Army knife of scientific models. Anything else from the related work that stands out? I think we covered the highlights, Nova. They took a lot from existing solutions and refined them, making significant advancements particularly in the way scientific data is handled and making efficient use of computational resources. Right. Echo. It feels like they're truly building bridges, both in data and in tech. It will be interesting to see how these advancements get utilized in practical applications. Absolutely. That wraps up our discussion on the related work section. Chatting through these advancements really emphasizes the collaborative nature of scientific progress, huh? For sure. Echo. It's all about standing on the shoulders of giants, but adding your own bit to reach even higher. All right. Folks, that brings us to the conclusion segment of today's episode. Nova, what an insightful journey we've had diving into this paper.
Absolutely. Echo. It's fascinating to see how the researchers behind Intern S1 Pro have pushed the envelope in AI for scientific discovery. So let's wrap up with some key takeaways and contributions from the paper. Shall we? Definitely. So first off, we have to highlight the monumental scale of this model. Intern S1 Pro boasts a whopping 1 trillion parameters, making it a powerhouse in both general and scientific domains. Yeah, and that's not just for show. This model's architectural innovations, like the expert expansion strategy and grouped routing, ensure efficient load balancing across devices. It also maintains training stability, which is crucial for large scale models. Right. And the researchers didn't stop there. They enhanced the model's scientific understanding through continued pre-training on 6 trillion tokens of high quality multimodal data. This involved a specialized caption pipeline tailored specifically for scientific imagery, which helped in producing precise and alignment focused captions.
Ah, the caption pipeline, such an essential tool. By generating these detailed captions, the model greatly improved its ability to interpret complex scientific visual content. This is a big step forward because it addresses the limitations of existing data sets in understanding scientific data visualizations. And then there's the rollout mechanism. This feature really stands out because it leverages prior scientific knowledge and helps generate complete reasoning chains. So instead of just recalling information, the model can perform logical deductions in complex scientific scenarios. Yeah, that's a game changer. Also, let's not forget the system prompt isolation strategy by creating mutually exclusive system level prefixes for scientific and general data. They manage to reduce data conflicts, improving both model stability and training effectiveness for one to source. It's pretty clear that all these innovations contribute to intern S1 pros, state-of-the-art performance across a wide range of scientific benchmarks.
The model displays robust reasoning capabilities and deep domain knowledge, which is incredibly impressive for one source. For sure, Echo, and looking ahead, the researchers plan to expand the model's capabilities even further into more specialized scientific domains. The goal is to accelerate scientific discovery, and it seems like they are well on their way. Exciting times indeed. Well, that's a wrap for today's episode. We hope you all enjoyed this deep dive into the fascinating world of AI-driven scientific discovery. Absolutely, folks. Thanks for tuning in. Don't forget to join us next time on the daily papercast for more intriguing papers and groundbreaking research. If you enjoyed today's episode, make sure to subscribe and leave us a review. Yeah. Feedback helps us grow and bring you even better content. Until next time, keep exploring the endless possibilities of AI. Bye for now. Bye, everyone. Stay curious and stay innovative.
More episodes
More from Daily Paper Cast

Seedance 2.0: Advancing Video Generation for World Complexity
Daily Paper Cast

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Age...
Daily Paper Cast

RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Tes...
Daily Paper Cast

SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Envir...
Daily Paper Cast