
索尼是日本一家全球知名的大型综合性跨国企业集团。总部设于日本东京都港区港南1-7-1。索尼是世界视听、电子游戏、通讯产品和信息技术等领域的先导者,是世界最早便携式数码产品的开创者,是世界最大的电子产品制造商之一、世界电子游戏业三大巨头之一、美国好莱坞六大电影公司之一。其旗下品牌有Xperia,Walkman,Sony Music,哥伦比亚电影公司,PlayStation等。曾有vaio旗下品牌,但在2014年2月6日,索尼剥离VAIO业务,Vaio品牌将由Japan Industrial Partners Inc接手运营。 2019年3月28日,索尼董事长平井一夫宣布退休,6月18日正式退休 ,交由公司原首席财务官吉田宪一郎担任。 2020年5月13日,索尼集团发布19~20财年财报,集团销售收入82599亿日元,实现营业利润8455亿日元。去除股权出售及上财年EMI业绩计入等一次性因素,营业利润同比增长1%。
Artificial intelligence (AI) systems now challenge or surpass human experts in many computer games1,2. Physical and real-time sports such as table tennis, however, remain a major open challenge because of their requirements for fast, precise and adversarial interactions near obstacles and at the edge of human reaction time3. Here we present Ace, to our knowledge the first real-world autonomous system competitive with elite human table tennis players. Ace addresses the challenges of physical real-time interaction through a new, high-speed perception system using event-based vision sensors4, and a new control system based on model-free reinforcement learning, as well as state-of-the-art high-speed robot hardware. Evaluated in matches against elite and professional players under official competition rules, Ace achieved several victories and demonstrated consistent returns of high-speed, high-spin shots. These results highlight the potential of physical AI agents to perform complex, real-time interactive tasks, suggesting broader applications in domains requiring fast, precise human-robot interaction.
Flow map models such as Consistency Models (CM) and Mean Flow (MF) enable few-step generation by learning the long jump of the ODE solution of diffusion models, yet training remains unstable, sensitive to hyperparameters, and costly. Initializing from a pre-trained diffusion model helps, but still requires converting infinitesimal steps into a long-jump map, leaving instability unresolved. We introduce *mid-training*, the first concept and practical method that inserts a lightweight intermediate stage between the (diffusion) pre-training and the final flow map training (i.e., post-training) for vision generation. Concretely, *Consistency Mid-Training* (CMT) is a compact and principled stage that trains a model to map points along a solver trajectory from a pre-trained model, starting from a prior sample, directly to the solver-generated clean sample. It yields a trajectory-consistent and stable initialization. This initializer outperforms random and diffusion-based baselines and enables fast, robust convergence without heuristics. Initializing post-training with CMT weights further simplifies flow map learning. Empirically, CMT achieves state-of-the-art two-step FIDs of 1.97 (CIFAR-10), 1.32 (ImageNet $64\times64$), and 1.84 (ImageNet $512\times512$), using up to $98$\% less training data and GPU time than CMs. On ImageNet $256\times256$, it attains 1-step FID 3.34 with $\sim50$\% less training than MF from scratch (FID 3.43). On MSCOCO T2I, CMT reaches the best FID with $\sim47$\% less training. This establishes CMT as a principled, efficient, and general framework for training flow map models. Code and models are available at https://github.com/sony/cmt.
Recent advances in generative foundational models, often termed "world models," have propelled interest in applying them to critical tasks like robotic planning and autonomous system training. For reliable deployment, these models must exhibit high physical fidelity, accurately simulating real-world dynamics. Existing physics-based video benchmarks, however, suffer from entanglement, where a single test simultaneously evaluates multiple physical laws and concepts, fundamentally limiting their diagnostic capability. We introduce WorldBench, a novel video-based benchmark specifically designed for concept-specific, disentangled evaluation, allowing us to rigorously isolate and assess understanding of a single physical concept or law at a time. To make WorldBench comprehensive, we design benchmarks at two different levels: 1) an evaluation of intuitive physical understanding with higher level concepts such as object permanence or scale/perspective, and 2) an evaluation of low-level physical constants and material properties such as friction coefficients or fluid viscosity, allowing to measure excatly how far from reality generated videos are. When SOTA video-based world models are evaluated on WorldBench, we find specific patterns of failure in particular physics concepts, with all tested models lacking the physical consistency required to generate reliable real-world interactions. Through its concept-specific evaluation, WorldBench offers a more nuanced and scalable framework for rigorously evaluating the physical reasoning capabilities of video generation and world models, paving the way for more robust and generalizable world-model-driven learning.
3D exoscopes have been increasingly adopted in otorhinolaryngology. However, conventional passive-polarized 3D displays (PPD) have limitations in vertical viewing angle, which can impair 3D visualization at nonfrontal angles, particularly for surgical assistants. Active-polarized 3D displays (APD) can overcome these limitations. This study aimed to compare the usability and performance of a conventional PPD and a prototype APD using the ORBEYE 4 K 3D exoscope system under viewing conditions unfavorable to the PPD. Twenty-four otorhinolaryngologists participated in the study. A prototype APD and a commercially available PPD were connected in parallel to the ORBEYE. The participants performed a procedural task simulating stapes surgery with the viewing position set 20° above the display center. The task performance was evaluated based on the number of successfully completed procedural tasks. A target-tracking test was performed before and after the procedural task to evaluate ocular fatigue by calculating the slope of the saccadic main sequence. The usability was assessed using a questionnaire. The APD scored higher than the PPD for all questionnaire items. The number of successful procedural tasks was significantly higher in the APD group. With the APD, there was no change in perceived ocular fatigue before and after the procedural task, whereas with the PPD, there was a tendency toward increased fatigue. The APD demonstrated superior usability and task performance compared to a conventional PPD, particularly under vertically displaced viewing conditions. APD may be particularly beneficial for assistants and surgeons working at various levels of the eye.
Music editing is an important step in music production, which has broad applications, including game development and film production. Most existing zero-shot text-guided editing methods rely on pretrained diffusion models by involving forward-backward diffusion processes. However, these methods often struggle to preserve the musical content. Additionally, text instructions alone usually fail to accurately describe the desired music. In this paper, we propose two music editing methods that improve the consistency between the original and edited music by leveraging score distillation. The first method, SteerMusic, is a coarse-grained zero-shot editing approach using delta denoising score. The second method, SteerMusic+, enables fine-grained personalized music editing by manipulating a concept token that represents a user-defined musical style. SteerMusic+ allows for the editing of music into user-defined musical styles that cannot be achieved by the text instructions alone. Experimental results show that our methods outperform existing approaches in preserving both music content consistency and editing fidelity. User studies further validate that our methods achieve superior music editing quality.