This data article describes the Surgical Procedure Phase, Keypoint, and Instrument Recognition (PhaKIR) dataset, a collection of eight complete laparoscopic cholecystectomy videos acquired during routine minimally invasive procedures at three German medical centers. The recordings were captured with different monocular endoscopic camera systems at 25 frames per second and a resolution of 1920 × 1080 pixels, with durations ranging from 28 to 58 min. Segments showing regions outside the abdominal cavity were removed for anonymization, and the corresponding cut indices are provided. The dataset offers unified annotations for three interrelated tasks on the same complete surgical sequences: surgical phase recognition, annotated for every frame at 25 fps (485,875 frames), instrument keypoint estimation (19,435 frames), and instrument instance segmentation (19,435 frames), the latter two annotated at one frame per second (every 25th frame). Surgical phases follow the seven-phase Cholec80 scheme, extended by an undefined label for transitional frames. The instrument annotations distinguish 19 instrument classes as well as individual instances of the same class, with keypoint annotations comprising two to four class-dependent points labeled with COCO-style visibility states. Segmentation and keypoint labels were created manually using the Computer Vision Annotation Tool (CVAT), whereas phase labels were derived from documented phase-transition timestamps, and all annotations underwent a multi-stage review by a medically trained team. The data are provided in open formats (MP4, CSV, JSON, PNG) together with a frame-extraction script under the CC BY-NC-SA license and are available through controlled access upon request via the Zenodo platform. Because the dataset combines procedural context, instrument pose, and pixel-accurate instance segmentations within full-length recordings from multiple institutions, it can be reused for single-task or multi-task model development, for temporally aware approaches that exploit motion continuity, and for cross-institutional generalization protocols such as leave-one-hospital-out evaluation. The dataset served as the training resource for the PhaKIR Challenge at the Endoscopic Vision (EndoVis) Challenge at MICCAI 2024.
Reliable recognition and localization of surgical instruments in endoscopic video recordings are foundational for a wide range of applications in computer- and robot-assisted minimally invasive surgery (RAMIS), including surgical training, skill assessment, and autonomous assistance. However, robust performance under real-world conditions remains a significant challenge. Incorporating surgical context - such as the current procedural phase - has emerged as a promising strategy to improve robustness and interpretability. To address these challenges, we organized the Surgical Procedure Phase, Keypoint, and Instrument Recognition (PhaKIR) sub-challenge as part of the Endoscopic Vision (EndoVis) challenge at MICCAI 2024. We introduced a novel, multi-center dataset comprising thirteen full-length laparoscopic cholecystectomy videos collected from three distinct medical institutions, with unified annotations for three interrelated tasks: surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation. Unlike existing datasets, ours enables joint investigation of instrument localization and procedural context within the same data while supporting the integration of temporal information across entire procedures. We report results and findings in accordance with the BIAS guidelines for biomedical image analysis challenges. The PhaKIR sub-challenge advances the field by providing a unique benchmark for developing temporally aware, context-driven methods in RAMIS and offers a high-quality resource to support future research in surgical scene understanding.
This recommendation provides a procedure for determining the carbonation depth on the surface of concrete by applying a pH indicator. This includes definitions of carbonation, carbonation depth and carbonation front, as well as descriptions of the different pH indicator solutions that can be used. Recommendations for testing laboratory-prepared specimens and those obtained from concrete structures are also given. This involves guidelines for sample preparation and/or extraction, CO2 exposure duration, carbonation depth determination and reporting of results. A section on data interpretation is also provided, as carbonation results are used for determining durability of concrete, as well as a criterion for materials selection or for carbon uptake calculations. The new Recommendation CPC-18R1 is intended to supersede the former RILEM recommendation CPC-18, particularly when prescribed as the preferred method for evaluating and reporting carbonation depths.
The utilisation of 3D printing processes in the fabrication of continuous fiber-reinforced composites confers a multitude of advantages, in particular flexible design based on structural requirements. In order to achieve greater flexibility, there is a necessity for 3D printing systems that allow for customisable material selection and fiber positioning. This paper presents the design of a robot-based 3D printing system that incorporates an in-situ impregnation line and flexibility regarding the machine code generation for fiber positioning. The development of the system enabled the attainment of an average fiber volume content of up to 37.12 E_1=24.7 GPa and strength of up to R_M1=0.51 GPa were determined.
We introduce a dataset and benchmark for cross-view urban traffic perception built from synchronized ego-centric bicycle videos and aerial drone videos recorded at real urban intersections. The benchmark targets two linked tasks: cross-view identity matching between street-view and drone-view object tracks, and ego-to-bird's-eye-view prediction using aerial supervision. In contrast to prior urban driving and V2X datasets, our benchmark provides identity-level alignment across radically different viewpoints together with standardized evaluation, annotation tooling, and baseline implementations. This setting is motivated by intersection-centric traffic analysis, where identity preservation, local interactions, and global spatial structure must be reasoned about jointly across views. We evaluate methods at both the track and frame levels, including cross-view ID precision/recall/IDF1, near–far breakdowns, temporal stability, and consistency metrics. We also provide baseline results for wedge-based cross-view matching and for three BEV prediction baselines: inverse perspective mapping, a MonoLayout-style learned baseline, and a regression baseline. The results show that the benchmark is feasible but challenging: cross-view matching achieves strong recall yet remains limited by over-assignment and temporal inconsistency, while ego-to-BEV prediction benefits from aerial supervision but remains far from saturated under lightweight monocular sensing. We hope that this benchmark will support future research on cross-view perception, urban scene alignment, and ego-to-global traffic understanding.