
We present an agile methodology based on the open source Chipyard framework used for designing and validating manufacturable and performant heterogeneous RISC-V SoCs within the constraints of 15-week semesters by classes composed primarily of undergraduate students. Chipyard integrates configurable, generator-based IP blocks and flows, including the modular VLSI flow, Hammer, developed over a decade of tapeouts in different technologies. Students iterate their custom RTL and AMS blocks through integration, verification, and place-and-route, then write full-stack applications, characterizing performance. One recent semester’s class chips in FinFET are described: COSMIC (FFT, convolution, DMA accelerators), MELLIS (sparse‑matrix, convolution, quantized transformer engines, near‑memory MAC), and SCμM‑V (low-power crystal-free transceiver, general-purpose AFE, on-chip power management, clock generation). For example, COSMIC reaches 1.25 GHz, accelerates compute 2-12× with energy savings, and runs live demos. New documentation and infrastructure, such as the new bring-up platform, Baremetal, make Chipyard even more accessible.
PEZY-SC4s is a fourth-generation many-core processor based on the multiple-instruction, multiple-data architecture, developed by PEZY Computing using TSMC 5nm process technology. It integrates 2,048 processor elements supporting 16,384 hardware threads and employs fine- and coarse-grained multithreading with a noncoherent hierarchical cache. This design achieves high energy efficiency while preserving general-purpose programmability, without relying on specialized tensor units or warp-based SIMT execution. The chip delivers a peak performance of 24.6 TFLOPS in double precision and 576 TFLOPS in BF16, backed by 3.2 TB/s of HBM3 memory bandwidth. Pre-silicon evaluation shows a double-precision matrix multiplication energy efficiency of 115 GFLOPS/W, a 2.2× improvement over the previous generation, and Smith–Waterman performance of 359 giga cell updates per second, a 3.9× improvement. A PyTorch-based software ecosystem enables deployment of major large language models, and a planned 90-node supercomputer system will deliver 8.9 PFLOPS in double precision.