
瑞萨科技是世界十大半导体芯片供应商之一,在很多诸如移动通信、汽车电子和PC/AV 等领域获得了全球最高市场份额。
As the wiring of LSI devices continues to shrink, the pitch of the bonding pads becomes smaller and smaller. This makes it more difficult to form connections using flip chip bumps and increases the cost of the interposer substrate. A direct fanning-out packaging technology, named as “SIRRIUS”, gives a solution for this problem. The SIRRIUS package consists of a Cu base plate, a fine-pad-pitch embedded LSI chip, built-up resin layers, and multi-layer fanning-out wiring. An experiment in which chips with 80-μm-pitch, four rows of pads are embedded clarifies that using a low-viscosity resin and optimizing its thickness are important for achieving fine-pitch interconnects with high reliability. By measuring the displacement of the mounted chips and laser vias, and improving their position accuracy, the feasibility of fabricating the SIRRIUS package with a lump method is demonstrated. A 31-mm-square, 900-pin-count SIRRIUS package with three-layer wiring, 15 μm wide, is successfully fabricated. Its total thickness is only 0.69 mm including the Cu plate. Evaluation of package-level and board-level reliability show no electrical failures or cracks after 600 and 1000 cycles, respectively. Our “SIRRIUS” package has following features; 1) small package thickness, 2) low cost process due to directly connecting the Cu plating to fine-pitch LSI pads, and 3) high heat dissipation, low package warpage and high reliability, due to the rigid Cu plate. Therefore, the SIRRIUS package is an attractive approach to achieving thin and highly reliable fine-pad-pitch system LSI packages for replacing FCBGA packages.
This paper presents the RX VPU, a newly developed Vector Processing Unit (VPU) designed as a coprocessor for the RX series of embedded microcontrollers. The RX VPU adopts a variable-length vector architecture and features a compact, low-power implementation suitable for embedded systems. It is specifically designed to accelerate CNN inference on embedded devices while maintaining low power consumption. The RX VPU provides instructions optimized for int8-quantized CNNs, which are widely used in embedded inference workloads, enabling efficient execution of compact CNN models. The architecture is particularly optimized for anomaly detection workloads using int8-quantized CNN models, where inner-product accumulation operations account for approximately 40% of all executed instructions. This work introduces a 4× widening reduction MAC instruction that accelerates these operations within a single instruction. Using this instruction, inference throughput on the Arm ML-Zoo anomaly detection model improves by 1.85×, achieving up to 30.0 inference/s. The coprocessor also achieves 1025.6 GFLOPS/W at 600 MHz, demonstrating that highly efficient ML inference can be achieved even under stringent power and area constraints.
This paper presents a topology-driven organic interposer design methodology for achieving RDL layer reduction in automotive Universal Chiplet Interconnect Express Advanced (UCIe-A) 2.5D chiplet packages while preserving signal integrity and manufacturability. By explicitly analyzing dominant crosstalk mechanisms in the bump-under region, the proposed methodology derives practical topology modifications based on via-row conversion and optimized ground (GND) shield placement under mass-producible design-rule constraints. The dominant crosstalk contributors are first identified and ranked using an X32 UCIe-A interface to enable clear attribution of inter-signal-layer, PAD-to-RDL, and multi-level crossing effects. Based on these insights, the topology-driven approach is extended to a full X64 interface representative of automotive applications. Cumulative distribution function (CDF)-based aggregate far-end crosstalk (FEXT) evaluation at 16 GHz, corresponding to the Nyquist frequency of a 32 Gbps NRZ signal, is employed to efficiently screen design variants and capture worst-case victims governing parallel-bus eye closure. Final signal integrity is validated through 32 Gbps eye-diagram simulations. Simulation results demonstrate that the proposed topology-driven design enables up to four-layer RDL reduction (e.g., from 10 to 6 layers) while preserving UCIe-A compliance, suppressing P95 and worst-case aggregate FEXT, and maintaining robust 32 Gbps eye openings. This work provides a scalable and manufacturable framework for cost-effective organic interposer design in high-reliability automotive UCIe-A 2.5D chiplet packages.
In edge AI inference, FPGA Overlay Accelerators offer superior performance-area balance but often lack accurate maximum frequency prediction and efficient Design Space Exploration (DSE) due to complex parameterization. To address these limitations, we propose a Bayesian optimization-based DSE framework using the Tree-Structured Parzen Estimator. We introduce a dependency-aware design space modeling (DAM) to prune the parameter space and a custom regression predictor (CREP) for frequency. To accommodate varying computational constraints, we present two implementations: a lightweight framework DEFA and an advanced framework DEFA-ML. DEFA saves computational cost of CREP and exploration runtime for rapid optimization. Conversely, DEFA-ML leverages a complex prediction model and a Large Language Model assistant for superior optimization quality. We validate the framework on the Altera FPGA AI Suite. For frequency prediction, DEFA provides fast and conservative predictions, while DEFA-ML achieves high linearity with a mean absolute percentage error of 4.21%. In throughput optimization, compared with the FPGA AI Suite baseline optimizer, DEFA improves throughput by 30.16% and DEFA-ML improves 35.12%. In area-throughput trade-off optimization, compared with the baseline, DEFA improves 39.05% and DEFA-ML improves 43.11% for single target model. For multiple target models, DEFA improves 40.76% and DEFA-ML improves 47.66%.
In this session, we present three talks from the IEEE standard (P3405) defining the standard multi-die interconnect test & repair methods describing various aspects of this upcoming standard. The first talk presents the standard test architecture & repair methods for the interconnects between the chiplets. The second talk presents a new description language to describe the underlying test & repair architecture of this standard, which can be used by CAD tools to automate the test vector generation & verifications of the interconnect tests between the chiplets. The third talk presents the fault modeling of the chiplet interconnects based on the defects that can occur in various packaging technologies like 3D, 2.5 D etc.