Large Reasoning Models (LRMs) have achieved remarkable progress thanks to Reinforcement Learning with Verifiable Rewards (RLVR) on Chain-of-Thoughts (CoTs). However, since long CoTs naturally contain trial and errors and mainstream RLVR approaches choose outcome-correct CoT trajectories for memorization, the redundant explorations in long CoTs are inevitably reinforced through RLVR, which results in the over-thinking issues of LRMs. Previous attempts to resolve the overthinking issue of LRMs mainly give more advantage to shorter trajectories, yet their learning signals are still outcome-based and cannot reduce the memorization of redundant explorations in long CoTs. Therefore, we propose ThoughtFold, a framework that leverages fine-grained preference learning to mitigate redundant explorations for efficient reasoning. ThoughtFold employs an introspective strategy to identify redundancy within each correct trajectory, which yields a spectrum of candidate sub-trajectories. Leveraging this spectrum, we introduce a masked preference optimization objective that explicitly penalizes redundant explorations and encourages the model to directly bridge essential reasoning segments, effectively folding its reasoning chains into a more concise path. Extensive experiments show that ThoughtFold significantly enhances efficiency. It reduces the token usage of DeepSeek-R1-Distill-Qwen-7B by approximately 56\% while maintaining state-of-the-art accuracy.
Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The training pipeline begins with scientific multimodal pre-training over rendered scientific documents, interleaved image-text data, and diverse scientific corpora. Starting from the pretrained checkpoint, we apply a unified post-training pipeline consisting of supervised fine-tuning, scalable multi-task reinforcement learning (RL), black- and white-box agentic RL, and on-policy distillation. This pipeline is supported by practical techniques that improve rollout and training stability and efficiency, including partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, robust multi-task optimization, and trace-aware experience assembly for agentic tasks. At the architecture level, Intern-S2-Preview-397B extends time series modelling from efficient long-sequence understanding to numerical forecasting, while Memory Decoder is studied as a separate memory-augmented path for rapid scientific specialization without modifying the frozen 397B backbone. Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings. The time series modules improve scientific signal understanding and forecasting on SciTS, while the separate Intern-MemDec-4B extension improves the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.
While the most fundamental pretraining paradigm typically trains modality-specific models on their respective datasets, the Platonic Representation Hypothesis that representations eventually align across modalities as data and model scale suggests an intriguing possibility: large language models (LLMs) could be pretrained on visual corpora to reach parity with text-pretrained models, thereby expanding data sources to break the text-scaling bottlenecks, and leveraging richer visual cues for more comprehensive corpus understanding. This paper makes the first attempt to demonstrate the feasibility of this implication by introducing Masked Autoregressive Pretraining for Learning language intelligencE (MAPLE), a novel visual pretraining paradigm for LLMs that leverages raw document images to improve language intelligence. MAPLE is universal to integrate masked auto-regressive models with various LLM backbones, where the LLMs are incentivized to generate latent hypotheses for the masked regions based on the unmasked regions. We verify MAPLE in the domain of math reasoning with multiple LLM backbones and show that MAPLE consistently surpasses text-only pretraining relatively by at most 40.2\% on average accuracy across four math reasoning benchmarks. Further analyses show that visually pretrained LLMs learn a shared latent space that aligns document visuals with text and exploits layout and structural cues, supporting visual pretraining as a feasible and scalable route to stronger language models.
Abstract Self-doped conducting polymers represent a distinct class of conjugated materials in which covalently tethered ionic functionalities enable intrinsic charge generation and stabilization. Unlike conventional externally doped systems, where weakly bound dopants often induce phase separation, diffusion, and long-term instability, self-doping establishes built-in charge compensation that enhances thermodynamic robustness and environmental tolerance. This review systematically examines how functional group engineering governs the electronic structure, charge transport, and interfacial properties of conducting polymers. Focusing on representative systems─polyaniline (PANI), poly(3,4-ethylenedioxythiophene) (PEDOT), and polypyrrole (PPy)─we elucidate the roles of sulfonate, carboxylate, quaternary ammonium, and phosphonate groups in modulating doping mechanisms, redox activity, and ion–electron coupling. By correlating molecular design with macroscopic performance across bioelectronics, energy storage, and optoelectronic applications, we establish a unified structure–property–function framework. Finally, we highlight unresolved challenges in doping efficiency, structural regulation, and device integration, providing design principles for next-generation stable and multifunctional conducting polymers.
Abstract Hydrogels, valued for their intrinsic biocompatibility and tunable mechanical properties, play a pivotal role in bioelectronics, biomedicine, and soft robotics. However, the intrinsic isotropy of conventional hydrogels limits their capacity to mimic the anisotropic architectures and directional functionalities of native tissues─hindering performance in applications requiring unidirectional force transmission, directional charge transport, or guided cell alignment. To address this, anisotropic hydrogels with spatially ordered architectures have emerged as transformative materials. This review systematically summarizes recent advances in the design and fabrication of anisotropic hydrogels, highlighting innovative strategies such as external-field-induced alignment and template-guided assembly, along with the integration of functional materials that enable precise molecular orientation and multiscale structural control. These engineered hydrogels exhibit programmable anisotropic responses─mechanical, electrical, and biological─unlocking new possibilities in flexible biosensing, intelligent actuation, targeted drug delivery, and tissue regeneration. Despite substantial progress, challenges remain in long-term structural stability, scalable manufacturing, biocompatibility of conductive fillers, and implementation of full-lifecycle design. Addressing these limitations through next-generation material innovations and fabrication technologies will be key to realizing the full potential of anisotropic hydrogels in advanced biointegrated systems.
Humans naturally excel at selective auditory attention, yet this ability is often impaired in individuals with hearing loss. Electroencephalography-based brain–computer interfaces (BCIs) can capture auditory task-evoked neural responses and assess attention states through brainwave analysis. However, traditional BCI systems rely on full-scalp wet electrodes and computation-heavy algorithms, limiting wearability and energy efficiency. Here we present a synergistic material–algorithm framework that combines an oxidant-free, room-temperature self-polymerization strategy to fabricate oligo(3,4-ethylenedioxythiophene)-based zwitterionic hydrogel electrodes with SDBformer, an ultra-lightweight, energy-efficient spiking Transformer algorithm. The hydrogel electrodes exhibit low on-skin impedance and stable in vivo electrocorticography for high-accuracy BCI control using steady-state visual evoked potentials, while SDBformer achieves robust auditory attention decoding using only eight temporal electroencephalography channels, matching the performance of full-scalp wet recordings. This integrated approach advances the development of practical, low-power wearable BCIs for neurotechnology applications in auditory attention and assistive hearing. An oxidant-free oligomer-based zwitterionic hydrogel electrode and an ultra-lightweight spiking transformer algorithm enable high-fidelity, low-power auditory attention decoding for practical wearable brain–computer interface applications.
The evolution of flexible and stretchable electronics demands organic field-effect transistors (OFETs) that harmonize high electrical performance with robust mechanical compliance. While material innovation and structural designs have advanced stretchability, a fundamental trade-off often persists between charge transport and elastic deformation. Side-chain engineering emerges as a precise molecular design strategy to transcend this compromise. By decorating conjugated polymer backbones with tailored functional groups, it is possible to intrinsically engineer intermolecular interactions, packing morphology, and energy dissipation mechanisms, thereby synergistically enhancing conductivity, stretchability, and even introducing other functions like self-healing. This review systematically examines this paradigm from a functional-group-centric perspective. We first categorize and elucidate the mechanisms of key side-chain groups-alkyl, hybrid, oligoether, fluoroalkyl, composite, and special functional groups-in regulating the mechanical-electronic property nexus. Subsequently, we critically analyze the application of these engineered materials in advanced functional devices, including skin-conformal sensors, stretchable displays and photodetectors, and neuromorphic computing circuits. Finally, we provide forward-looking perspectives on the challenges and opportunities in designing next-generation side-chains for multifunctional, reliable, and commercially viable stretchable electronics.
We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond stronger reasoning and image-text understanding capabilities, its intelligence is augmented with advanced agent capabilities. Simultaneously, its scientific expertise has been vastly expanded to master over 100 specialized tasks across critical science fields, including chemistry, materials, life sciences, and earth sciences. Achieving this massive scale is made possible by the robust infrastructure support of XTuner and LMDeploy, which facilitates highly efficient Reinforcement Learning (RL) training at the 1-trillion parameter level while ensuring strict precision consistency between training and inference. By seamlessly integrating these advancements, Intern-S1-Pro further fortifies the fusion of general and specialized intelligence, working as a Specializable Generalist, demonstrating its position in the top tier of open-source models for general capabilities, while outperforming proprietary models in the depth of specialized scientific tasks.
Hydrogels, owing to their exceptional flexibility, stretchability, and biocompatibility, have emerged as ideal substrates for wearable bioelectronics, biomedical devices, and related applications. Nevertheless, conventional hydrogels are prone to mechanical damage under prolonged external stress and are susceptible to dehydration, swelling, and contamination under extreme environmental conditions, compromising their stability and practical performance. To address these limitations, various strategies have been developed to enhance hydrogel stability, including structural optimization and the incorporation of functionalized or modified materials. This review systematically summarizes key approaches for improving hydrogel mechanical robustness, self-healing ability, water retention, freeze resistance, antiswelling performance, and antimicrobial activity. The underlying mechanisms, advantages, limitations, and application scenarios of these strategies are critically analyzed. Finally, future perspectives on the design and development of hydrogels with superior stability, as well as their potential applications in advanced biomedical and bioelectronic fields, are discussed.
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provide limited disciplinary coverage and often rely on final-answer correctness or coarse judgments, leaving the validity of the reasoning process inadequately assessed. To bridge this gap, we introduce AdvancedMathBench, a benchmark suite designed to evaluate advanced mathematical reasoning capabilities. Its core proof-generation benchmark, ProverBench, contains 296 problems spanning undergraduate and doctoral qualifying-exam levels. To provide reliable evaluation of the proofs, we develop a dedicated automatic verification pipeline trained on large-scale expert annotations to produce both correctness verdicts and fine-grained assessments of proof errors, which exhibits strong agreement with human experts on held-out proof trajectories. We further introduce VerifierBench, consisting of 888 model-generated proof trajectories paired with expert ground truth, to evaluate whether models can correctly judge proof validity and provide sound verification rationales. Experiments show that AdvancedMathBench remains challenging for frontier models. On proof generation, the best-performing model, GPT-5.5-xhigh, achieves only 75.8 and 66.1 on the UGD and QE splits, respectively, indicating substantial room for improvement on advanced mathematical proof construction. On proof verification, the best model attains a Balanced F1 of only 65.1, and models generally exhibit low true negative rates, suggesting that critical error detection remains a major bottleneck.
Skin conductance (SC) is a core physiological signal that can directly characterize human sympathetic nerve activity, emotional fluctuations, and cognitive states, creating an urgent demand for long-term, accurate, and non-invasive SC monitoring technology in the fields of intelligent medical treatment and human–computer interaction. Traditional measurement methods based on rigid electrodes suffer from poor skin interface matching and severe motion artifacts, which fail to meet the requirements of real-time monitoring in dynamic scenarios. Starting from the physiological basis of SC, this review illustrates the core mechanism of sweat gland activity regulated by sympathetic nerves, and elaborates in detail the measurement principle and key technologies based on Ohm’s law. At the device design level, this paper focuses on the key design points of ionic conductors, electronic conductors, and composite materials, analyzes the optimization effect of micro-nanostructure innovation on device performance, and summarizes the typical applications of the technology. In view of the core challenges faced by current flexible electrodes, this paper points out that future research will focus on the collaborative breakthrough of environmental adaptability and personalized matching, providing a systematic theoretical reference for the research, development and application of SC measurement technology.
Length generalization, the ability to solve problems of longer sequences than those observed during training, poses a core challenge of Transformer-based large language models (LLM). Although existing studies have predominantly focused on data-driven approaches for arithmetic operations and symbolic manipulation tasks, these approaches tend to be task-specific with limited overall performance. To pursue a more general solution, this paper focuses on a broader case of reasoning problems that are computable, i.e., problems that algorithms can solve, thus can be solved by the Turing Machine. From this perspective, this paper proposes Turing MAchine Imitation Learning (TAIL) to improve the length generalization ability of LLMs. TAIL synthesizes chain-of-thoughts (CoT) data that imitate the execution process of a Turing Machine by computer programs, which linearly expands the reasoning steps into atomic states to alleviate shortcut learning and explicit memory fetch mechanism to reduce the difficulties of dynamic and long-range data access in elementary operations. To validate the reliability and universality of TAIL, we construct a challenging synthetic dataset covering 8 classes of algorithms and 18 tasks. Without bells and whistles, TAIL significantly improves the length generalization ability as well as the performance of Qwen2.5-7B on various tasks using only synthetic data, surpassing previous methods and DeepSeek-R1. The experimental results reveal that the key concepts in the Turing Machine, instead of the thinking styles, are indispensable for TAIL for length generalization, through which the model exhibits read-and-write behaviors consistent with the properties of the Turing Machine in their attention layers. This work provides a promising direction for future research in the learning of LLM reasoning from synthetic data.
While current Multimodal Large Language Models (MLLMs) have demonstrated proficiency in reasoning tasks such as mathematics and logic, their capacity for long-chain reflective reasoning, a prerequisite for solving complex real-world problems, remains largely underexplored. In this work, we first conduct an extensive empirical investigation to evaluate this capability. Leveraging a carefully designed data synthesis engine, we construct MM-HELIX, a multimodal benchmark consisting 1260 samples of 42 challenging synthetic tasks that require iterative thinking and backtracking. Empirical results on this benchmark reveal that existing MLLMs exhibit significant performance deficits in long-chain reflective reasoning. To address this limitation, we generate post-training data and further explore learning paradigms for exploiting such data. We first develop the Step-Elicited Response Generation pipeline to create MM-HELIX-100K, a large-scale dataset of 100k high-quality, reflective reasoning traces for instruction-tuning stage. Given that standard Reinforcement Learning fails on complex tasks due to sparse reward signals and catastrophic forgetting after Supervised Fine-Tuning, we propose Adaptive Hybrid Policy Optimization (AHPO), a novel training strategy that dynamically unifies offline supervision and online optimization into a single stage. This strategy enables the model to learn from expert data when rewards are sparse and conduct independent exploration once proficient. When applied to the Qwen2.5-VL-7B baseline, our method achieves a +18.6\% accuracy improvement on MM-HELIX benchmark and demonstrates strong generalization with a +5.7\% average performance gain on general mathematic and logic tasks. Our work demonstrate that reflective reasoning in MLLMs can be effectively learned and generalized, paving the way for developing more capable MLLMs.
Flexible ferroelectric pressure sensors have emerged as a key technological platform for next-generation flexible electronics owing to their self-powered operation, high sensitivity, and excellent mechanical compliance. This Review systematically summarizes recent advances in the field, with particular emphasis on the ongoing transition from single-component optimization to multidimensional synergistic design across ferroelectric materials, functional substrates, flexible electrodes, and device architectures. We first outline the fundamental working principles of representative ferroelectric systems and then discuss major strategies for performance enhancement, including molecular design of organic ferroelectrics, interface engineering in polymer–ceramic composites, flexible integration of high-performance inorganic materials, and microstructure engineering for stress modulation and signal amplification. Particular attention is paid to hierarchical microstructures that help alleviate the long-standing trade-off between sensitivity and linear working range. Representative applications in wearable health monitoring, electronic skin, and human–machine interaction are further highlighted. Finally, current challenges and future opportunities are discussed from the perspectives of material reliability, scalable manufacturing, multifunctional integration, and the convergence of multimodal sensing with artificial intelligence.
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that cannot be faithfully or completely captured by text alone. Yet current pretraining approaches discard these visual cues by converting visually rich sources, such as documents and web pages, into plain text for learning language intelligence. This paper challenges the default assumption that language models must be trained on text-only representations and shows that Visual Pretraining is a scalable learner for foundation model intelligence. To this end, we conduct a systematic study of unsupervised visual pretraining paradigms that directly leverage visual documents without text extraction. Across multiple backbones and benchmarks, visual pretraining on the same underlying corpora consistently outperforms text-only pretraining, offering an efficient pathway to scalable language intelligence.
Large Reasoning Models (LRMs) have expanded the mathematical reasoning frontier through Chain-of-Thought (CoT) techniques and Reinforcement Learning with Verifiable Rewards (RLVR), capable of solving AIME-level problems. However, the performance of LRMs is heavily dependent on the extended reasoning context length. For solving ultra-hard problems like those in the International Mathematical Olympiad (IMO), the required reasoning complexity surpasses the space that an LRM can explore in a single round. Previous works attempt to extend the reasoning context of LRMs but remain prompt-based and built upon proprietary models, lacking systematic structures and training pipelines. Therefore, this paper introduces Intern-S1-MO, a long-horizon math agent that conducts multi-round hierarchical reasoning, composed of an LRM-based multi-agent system including reasoning, summary, and verification. By maintaining a compact memory in the form of lemmas, Intern-S1-MO can more freely explore the lemma-rich reasoning spaces in multiple reasoning stages, thereby breaking through the context constraints for IMO-level math problems. Furthermore, we propose OREAL-H, an RL framework for training the LRM using the online explored trajectories to simultaneously bootstrap the reasoning ability of LRM and elevate the overall performance of Intern-S1-MO. Experiments show that Intern-S1-MO can obtain 26 out of 35 points on the non-geometry problems of IMO2025, matching the performance of silver medalists. It also surpasses the current advanced LRMs on inference benchmarks such as HMMT2025, AIME2025, and CNMO2025. In addition, our agent officially participates in CMO2025 and achieves a score of 102/126 under the judgment of human experts, reaching the gold medal level.
Large Language Models (LLMs) are now widely used across many domains. With their rapid development, Reinforcement Learning with Verifiable Rewards (RLVR) has surged in recent months to enhance their reasoning and understanding abilities. However, its complex data flows and diverse tasks pose substantial challenges to RL training systems, and there is limited understanding of RLVR from a system perspective. To thoroughly understand the system challenges introduced by RLVR, we present a characterization study of RLVR tasks in our LLM deployment. Specifically, we investigate the distribution and variation trends of workloads across different RL tasks across training steps. We identify issues such as GPU idling caused by skewed sequence length distribution, inefficient parallel strategies in dynamically varying workloads, inefficient data management mechanisms, and load imbalance. We describe our observations and call for further investigation into the remaining open challenges. Furthermore, we propose PolyTrace benchmark suite to conduct evaluation with realistic workloads, and a practical use case validates that PolyTrace benchmark suite exhibits 94.7