Purpose: Large language models (LLMs) may reduce administrative workload in radiology by automating structured diagnostic extraction from text reports. This study evaluates the accuracy of ChatGPT-4.0 when extracting correct diagnoses from musculoskeletal (MSK) radiology text reports, and compares its performance with that of experienced human readers, using cluster-adjusted and consensus-level analyses. Materials and Methods: Twenty-three multimodal MSK imaging cases (X-ray, ultrasound, CT, and MRI) were analysed. Ten human readers and ChatGPT-4.0 (10 independent iterations) provided primary (1st) and secondary (2nd) diagnoses from six predefined options. We analysed data at the individual-reader level using cluster-adjusted generalised estimating equations (GEE) and at the case level using majority consensus with exact McNemar testing. Within-case (α_case) and within-reader (α_reader) correlations and design effects were calculated to assess clustering and implications for sample size. Results: For 1st diagnoses, AI accuracy was 0.957 (95%–CI 0.922–0.976) versus 0.865 (95%–CI 0.815–0.903) for human readers (absolute difference −0.091; OR 3.43, 95%–CI 1.07–11.02; p = 0.038). Within-case correlation (α case = 0.247) exceeded within-reader correlation (α reader ≈ 0); this resulted in a design effect of 5.7 and an effective sample size of 80.7. At the consensus level, discordance occurred in 2/23 cases (8.7%), with no significant difference between methods (McNemar p = 1.00). When 1st and 2nd diagnoses were combined, both systems achieved 23/23 correct consensus classifications. Interrater reliability between AI and human classifications was almost perfect (Gwet’s AC1 = 0.836–0.927). Conclusions and Key points: In this structured MSK text-report setting, ChatGPT-4.0 achieved diagnostic accuracy comparable to that of experienced radiologists, with modest individual-reader advantages that disappeared under consensus aggregation. Clustering analysis indicates that variability is primarily case-driven, suggesting that future validation studies will benefit more from expanding case numbers than reader numbers. Our data suggest that large performance divergences between AI and human consensus are unlikely in similar structured diagnostic contexts.
Legitimate gatekeeping in science that is based on the standards constitutive of science presupposes a distinction between science and non-science and thus a solution to the demarcation problem. However, the latter has often been considered as unsolved, if not unsolvable. This paper aims to trace the consequences of discussions about the demarcation problem for gatekeeping. I argue that an approach that relativizes scientific status to time, research field, and level allows for a meaningful demarcation and thus gatekeeping. Particular emphasis is put on the last point, i.e. the level comprising the units to which we wish to attribute scientific status, e.g. ideas, experiments, or publications. I conduct an empirical analysis to identify the kinds of things that are called “scientific”, “unscientific”, or “pseudoscientific” in ordinary talk. To classify the results, I propose a systematic coding scheme based on the idea that science is a practice that leads from starting points to results. I consider a few units, such as ideas and experiments, and discuss the criteria to be applied for these units and the consequences for gatekeeping. Such a unit-sensitive account has the advantage of being more lenient concerning the starting points of research and more stringent regarding the results of scientific inquiry.
Novel artificial intelligence tools have the potential to significantly enhance productivity in medicine, while also maintaining or even improving treatment quality. In this study, we aimed to evaluate the current capability of ChatGPT-4.0 to accurately interpret multimodal musculoskeletal tumor cases.We created 25 cases, each containing images from X-ray, computed tomography, magnetic resonance imaging, or scintigraphy. ChatGPT-4.0 was tasked with classifying each case using a six-option, two-choice question, where both a primary and a secondary diagnosis were allowed. For performance evaluation, human raters also assessed the same cases.When only the primary diagnosis was taken into account, the accuracy of human raters was greater than that of ChatGPT-4.0 by a factor of nearly 2 (87% vs. 44%). However, in a setting that also considered secondary diagnoses, the performance gap shrank substantially (accuracy: 94% vs. 71%). Power analysis relying on Cohen's w confirmed the adequacy of the sample set size (n: 25).The tested artificial intelligence tool demonstrated lower performance than human raters. Considering factors such as speed, constant availability, and potential future improvements, it appears plausible that artificial intelligence tools could serve as valuable assistance systems for doctors in future clinical settings. · ChatGPT-4.0 classifies musculoskeletal cases using multimodal imaging inputs.. · Human raters outperform AI in primary diagnosis accuracy by a factor of nearly two.. · Including secondary diagnoses improves AI performance and narrows the gap.. · AI demonstrates potential as an assistive tool in future radiological workflows.. · Power analysis confirms robustness of study findings with the current sample size.. · Bosbach WA, Schoeni L, Beisbart C et al. Evaluating the Diagnostic Accuracy of ChatGPT-4.0 for Classifying Multimodal Musculoskeletal Masses: A Comparative Study with Human Raters. Rofo 2025; DOI 10.1055/a-2594-7085.
Despite their successes at prediction and classification, deep neural networks (DNNs) are often claimed to fail when it comes to providing any understanding of real-world phenomena. However, recently, some authors have argued that DNNs can provide such understanding. To resolve this controversy, I first examine under which conditions DNNs provide humans with explanatory understanding in a clearly defined sense that refers to a simple setting. I adopt a systematic approach that draws on theories of explanation and explanatory understanding, but avoid dependence on any specific account by developing broad conditions of explanatory understanding that leave space for filling in the details in several alternative ways. I argue that the conditions are difficult to satisfy however these details are filled in. The main problem is that, to provide explanatory understanding in the sense I have defined, a DNN has to contain an explanation, and scientists typically do not know whether it does. Accordingly, they cannot feel committed to the explanation or use it, which means that other conditions of explanatory understanding are not satisfied. Still, in some attenuated senses, the conditions can be fulfilled. To complete my conciliatory project, I further show that my results so far are compatible with using DNNs to infer explanatorily relevant information in a thorough investigation. This is what the more optimistic literature on DNNs has focused on. In sum, then, the significance of DNNs for understanding real-world systems depends on what it means to say that they provide understanding, and on how humans use them.
The goal of this paper is to re-assess reflective equilibrium (“RE”). We ask whether there is a conception of RE that can be defended against the various objections that have been raised against RE in the literature. To answer this question, we provide a systematic overview of the main objections, and for each objection, we investigate why it looks plausible, on what standard or expectation it is based, how it can be answered and which features RE must have to meet the objection. We find that there is a conception of RE that promises to withstand all objections. However, this conception has some features that may be unexpected: it aims at a justification that is tailored to understanding and it is neither tied to intuitions nor does it imply coherentism. We conclude by pointing out a cluster of questions we think RE theorists should pay more attention to.
In contemporary computer simulations, particles attract each other and form clusters, cells interact, and agents communicate with one another. This is at least how computer simulations are commonly described. But how can we make sense of such talk? One answer is that the particles, cells, and agents inside simulations are digital artifacts, and thus real objects. In this paper, I cast doubt on this realist position by raising the question: To what objects does a simulation give rise, if it does so at all? I tentatively suggest answers that try to determine what objects arise in computer simulations at both a fundamental and a higher level. However, closer analysis reveals that we run into underdetermination problems. Even for a single run of a simulation program, the acknowledged facts about the simulations often do not suffice to determine the objects and their properties uniquely. I suggest that, from an ontological perspective, it is more attractive to say that computer simulations involve computations, the results of which may then be interpreted in various ways. To make sense of talk about objects in simulations it is more attractive to develop a fictionalist account, or so I argue.
Physik in unserer ZeitVolume 55, Issue 2 p. 100-100 Magazin Unbestimmt und relativ? Das Weltbild der modernen Physik Claus Beisbart, Claus Beisbart BernSearch for more papers by this author Claus Beisbart, Claus Beisbart BernSearch for more papers by this author First published: 01 March 2024 https://doi.org/10.1002/piuz.202470212Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onEmailFacebookTwitterLinkedInRedditWechat No abstract is available for this article. Volume55, Issue2March 2024Pages 100-100 RelatedInformation
Motive: Documentation and administration, unpleasant necessities, take a substantial part of the working time in the subspecialty of interventional radiology. With increasing future demand for clinical radiology predicted, time savings from use of text drafting technologies could be a valuable contribution towards our field. Method: Three cases of peripherally inserted central catheter (PICC) line insertion were defined for the present study. The current version of ChatGPT was tasked with drafting reports, following the Radiological Society of North America (RSNA) template. Key results: Score card evaluation by human radiologists indicates that time savings in documentation / administration can be expected without loss of quality from using ChatGPT. Further, automatically generated texts were not assessed to be clearly identifiable as AI-produced. Conclusions: Patients, doctors, and hospital administrators would welcome a reduction of the time that interventional radiologists need for documentation and administration these days. If AI-tools as tested in the present study are brought into clinical application, questions about trust into those systems eg with regard to medical complications will have to be addressed.
The amount of acquired radiology imaging studies grows worldwide at a rapid pace. Novel information technology tools for radiologists promise an increase of reporting quality and as well quantity at the same time. Automated text report drafting is one branch of this development. We defined for the present study in total 9 cases of distal radius fracture. Command files structured according to a template of the Radiological Society of North America (RSNA) and to Arbeitsgemeinschaft Osteosynthese (AO) classifiers were given as input to the natural language processing tool ChatGPT. ChatGPT was tasked with drafting an appropriate radiology report. A parameter study (n = 5 iterations) was performed. An overall high appraisal of ChatGPT radiology report quality was obtained in a score card based assessment. ChatGPT demonstrates the capability to adjust output files in response to minor changes in input command files. Existing shortcomings were found in technical terminology and medical interpretation of findings. Text drafting tools might well support work of radiologists in the future. They would allow a radiologist to focus time on the observation of image details and patient pathology. ChatGPT can be considered a substantial step forward towards that aim.
Since the advent of digital computers around 1950, the method of computer simulation (simulation, for short) has enriched the repertoire of scientific methods. A computer simulation traces the dynamical behaviour of a target system by evaluating a (possibly partial and approximate) solution to a model. Philosophers of science have clarified the concept of computer simulation and its subcategories, analysed the justification of simulation results, explained the power and limitations of simulations, and explored their broader significance for science.
In a computer simulation, a digital computer is used to trace the time evolution of a system, e.g., of the atmosphere of the Earth. Using a model, the computer calculates the values of variables such as air pressure for a series of times and thus obtains state descriptions of the system for those times. The outputs of computer simulations are often visualized using animations. If all goes well in the simulation, humans can learn from the outputs how the system under consideration evolves with time. Computer simulations became possible with the advent of the digital computer in the 1940s. They contrast with analog simulations, which do without a digital computer and are not considered in this entry. Among the first computer simulations were programs that were intended to trace the explosion of nuclear weapons and the evolution of the weather. Nowadays, the use of computer simulations is widespread, in particular in education and in research in the natural and social sciences. But not every use of the computer qualifies as computer simulation; for instance, the classification of images using neural networks does not count as computer simulation because no time evolution is traced. In philosophy, computer simulation is mainly discussed as a scientific practice or method. Accordingly, it is mostly philosophers of science who study computer simulation. Their focus has been on the epistemology of computer simulation: They study how computer simulations are embedded in, and change, the workings of science. The most important questions are: (1) How are computer simulations used in different disciplines and what kinds of tasks do they fulfill? (2) How are simulation results justified? (3) How can we explain how computer simulations achieve their tasks? Since answers to this question often relate computer simulation to other methods, e.g., experimentation, they also tell us what kind of method computer simulation is. (4) To what extent are computer simulations novel in science and what are the consequences for our philosophical picture of science? The last question is pressing because well-established positions in philosophy of science, e.g., falsificationism, Bayesianism, and Kuhn’s position, have been developed without reference to computer simulation. An additional important topic that has emerged in the philosophical discussion of simulations is their black box character. This trait is often labeled “epistemic opacity” and seems relevant for the question of how computer simulations can provide understanding. The current wave of interest in computer simulations started in the 1990s, although there were a few philosophical publications on the theme before that time. This bibliography covers this wave without going further back to the past. The focus is entirely on philosophical appraisals of simulations within science. The use of simulations as a tool within philosophy is largely bracketed.
Motive Documentation and administration, unpleasant necessities, take a substantial part of the working time in the subspecialty of interventional radiology. With increasing future demand for clinical radiology predicted, time savings from use of text drafting technologies could be a valuable contribution towards our field. Method Three cases of peripherally inserted central catheter (PICC) line insertion were defined for the present study. The current version of ChatGPT was tasked with drafting reports, following the Radiological Society of North America (RSNA) template. Key results Score card evaluation by human radiologists indicates that time savings in documentation / administration can be expected without loss of quality from using ChatGPT. Further, automatically generated texts were not assessed to be clearly identifiable as AI-produced. Conclusions Patients, doctors, and hospital administrators would welcome a reduction of the time that interventional radiologists need for documentation and administration these days. If AI-tools as tested in the present study are brought into clinical application, questions about trust into those systems eg with regard to medical complications will have to be addressed. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement Acknowledgements and funding: The authors wish to thank for all the useful discussions leading to this manuscript. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: N/A I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Online supplement: study raw data deposited under doi.org/10.5281/zenodo.8140755
This paper details the content and structure of the folk concept of life, and discusses its relevance for scientific research on life. In four empirical studies, we investigate which features of life are considered salient, universal, central, and necessary. Functionings, such as nutrition and reproduction, but not material composition, turn out to be salient features commonly associated with living beings (Study 1). By contrast, being made of cells is considered a universal feature of living species (Study 2), a central aspect of life (Study 3), and our best candidate for being necessary for life (Study 4). These results are best explained by the hypothesis that people take life to be a natural kind subject to scientific scrutiny.
In his Critique of Pure Reason, in the chapter on the antinomy of pure reason, Kant not only argues that aprioristic cosmology is doomed to failure; he also implies that empirical knowledge about the universe is impossible. Today, such a negative verdict about the possibility of cosmological knowledge seems implausible because physical cosmology has made substantial progress. In particular, the spatiotemporal extension of the universe now seems a matter of empirical investigation in which models figure centrally. But I think it is worth considering the possibility that Kant got something right and that he offers insights that can help us to better understand problems in present-day cosmology. In this article, I explore a striking coincidence: according to both Kant and current wisdom, cosmology faces a serious underdetermination problem regarding the spatiotemporal extension of the world. As a closer analysis reveals, however, Kant and modern cosmology differ on the reasons why underdetermination arises. In current cosmology, underdetermination follows from laws that are knowable a posteriori, and not only from the very idea of cosmological knowledge, as Kant would have it. This suggests that the current underdetermination problem is not fully a Kantian one.
In computer science, there are efforts to make machine learning more interpretable or explainable, and thus to better understand the underlying models, algorithms, and their behavior. But what exactly is interpretability, and how can it be achieved? Such questions lead into philosophical waters because their answers depend on what explanation and understanding are-and thus on issues that have been central to the philosophy of science. In this paper, we review the recent philosophical literature on interpretability. We propose a systematization in terms of four tasks for philosophers: (i) clarify the notion of interpretability, (ii) explain the value of interpretability, (iii) provide frameworks to think about interpretability, and (iv) explore important features of it to adjust our expectations about it.
Some machine learning models, in particular deep neural networks (DNNs), are not very well understood; nevertheless, they are frequently used in science. Does this lack of understanding pose a problem for using DNNs to understand empirical phenomena? Emily Sullivan has recently argued that understanding with DNNs is not limited by our lack of understanding of DNNs themselves. In the present paper, we will argue, contra Sullivan, that our current lack of understanding of DNNs does limit our ability to understand with DNNs. Sullivan's claim hinges on which notion of understanding is at play. If we employ a weak notion of understanding, then her claim is tenable, but rather weak. If, however, we employ a strong notion of understanding, particularly explanatory understanding, then her claim is not tenable.
Reflective equilibrium (RE) is often regarded as a powerful method in ethics, logic, and even philosophy in general. Despite this popularity, characterizations of the method have been fairly vague and unspecific so far. It thus may be doubted whether RE is more than a jumble of appealing but ultimately sketchy ideas that cannot be spelled out consistently. In this paper, we dispel such doubts by devising a formal model of RE. The model contains as components the agent’s commitments and a theory that tries to systematize the commitments. It yields a precise picture of how the commitments and the theory are adjusted to each other. The model differentiates between equilibrium as a target state and the dynamic equilibration process. First solutions to the model, obtained by computer simulation, show that the method allows for consistent specification and that the model’s implications are plausible in view of expectations on RE. In particular, the mutual adjustment of commitments and theory can improve one’s commitments, as proponents of RE have suggested. We argue that our model is fruitful not only because it points to issues that need to be dealt with for a better understanding of RE, but also because it provides the means to address these issues.
Does complexity make multilingualism special? Since there is no unequivocal notion of complexity on which researchers agree, several characteristics that have been considered crucial for complexity are brought to bear on multilingualism. While multilingualism is fairly complex in some senses, for instance, because it requires that many variables be studied, it is less clear whether multilingualism becomes special in this way. The most salient possible way in which multilingualism might be special due to its complexity is that qualitatively new features emerge as we move from mono- or bilingualism to multilingualism. While research in mathematics and physics has shown examples in which novel features emerge at a certain level, related results are not easily transferred to multilingualism. This is shown by analyzing the application of dynamic systems theory to multilingualism by Herdina and Jessner. It is ultimately a matter of empirical research whether there are suitable novel features that make multilingualism special.