Credit card fraud constitutes a core component of the contemporary cybercrime economy, in which dark web carding forums play a pivotal role in coordinating, commoditising, and disseminating illicit activities. While prior research has primarily focused on transaction-level fraud detection, comparatively limited attention has been devoted to the systematic analysis of the social and organisational ecosystems within which these practices are enacted. This study addresses this gap by proposing and validating a domain-specific taxonomy for the automated classification of content in P2P carding forums. To this end, we adopt an iterative, data-driven methodology that integrates large language models (LLMs), lexical co-occurrence analysis, and semantic network analysis. Using a corpus of 3260 posts, we define and operationalise a taxonomy structured around four predicates: activity context, actor role, products and services, and technical tools, supported by a locally deployed LLM (Llama 4 Scout). A human-annotated subset was additionally used to evaluate inter-annotator agreement and standard classification metrics, complementing the coverage-based assessment and enabling comparison against a keyword-based baseline. Evaluation was further strengthened through manual benchmarking, confidence intervals, sensitivity analysis of key pipeline components, and comparison with alternative open-weight models. The results indicate that the proposed taxonomy achieves broad corpus-level representational coverage, with at least one semantic dimension identified in 98.71% of posts. However, coverage is uneven across predicates: activity-context is highly explicit, whereas actor-role and product-service show only moderate coverage and technique-tool remains substantially underrepresented and ambiguous. Overall, the findings show that combining domain-specific taxonomies with LLM-assisted classification and network analysis offers a robust framework for understanding and monitoring carding ecosystems in the dark web.
Large-scale content analysis is increasingly limited by the absence of observable ground truth or gold-standard labels, as creating such benchmarks through extensive human coding becomes impractical for massive datasets due to high time, cost, and consistency challenges. To overcome this barrier, we introduce the AI-CROWD protocol, which approximates ground truth by leveraging the collective outputs of an ensemble of large language models (LLMs). Rather than asserting that the resulting labels are true ground truth, the protocol generates a consensus-based approximation derived from convergent and divergent inferences across multiple models. By aggregating outputs via majority voting and interrogating agreement/disagreement patterns with diagnostic metrics, AI-CROWD identifies high-confidence classifications while flagging potential ambiguity or model-specific biases.
This study evaluates the classification performance of four GPT-based models (GPT-4.1, GPT-4.1-mini, GPT-4.1-nano, and o4-mini) under zero-shot prompting conditions on the complete, multilingual CoDA dataset of Dark Web content, comprising 10 illicit activity categories. The models GPT-4.1, GPT-4.1-mini, and o4-mini achieve a weighted F1 score of 0.885, surpassing prior zero-shot baselines on this dataset. Stability analysis using TARa@10 demonstrates high output consistency for GPT-4.1 (0.964) and GPT-4.1-mini (0.970), indicating their reliability for operational use. Multilingual evaluation reveals only a modest English vs. non-English performance gap for GPT-4.1 (0.031), while other models perform comparably across languages. The strongest results appear in Drugs, Gambling, and Porn (F1 > 0.9), whereas lower scores are observed in ambiguous or overlapping categories like Violence (F1 <= 0.76) or Crypto (F1 <= 0.84). A qualitative review of misclassifications suggests that some model predictions align with reasonable semantic interpretations, potentially highlighting annotation inconsistencies. This work establishes a performance baseline for GPT-based models in zero-shot classification of multilingual Dark Web content and underscores the importance of clear category definitions for effective deployment.
Dark web question-and-answer (Q&A) forums hosted on the Tor network serve as critical hubs for cyber threats and social interactions, yet their diverse content challenges traditional classification methods. This study leverages large language models (LLMs) to classify and extend the MISP dark web taxonomy for analyzing 2,055 substantive posts from three high-traffic dark web Q&A forums, scraped between July and November 2024. A multi-stage pipeline was employed: initial classification with the Mistral 7B model identified dominant topics (e.g., "finance-crypto," 10.80%) and motivations (e.g., "forum," 20.54%), revealing high ambiguity (40.34% for topics, 10.37% for motivations). HDBSCAN clustering of ambiguous posts uncovered novel themes, including "mental-health" and "confessions-and-personal-secrets," prompting the extension of the MISP taxonomy with seven new topic categories and one motivation category. Reclassification reduced ambiguity to 2.95% for topics and 1.51% for motivations, enhancing the taxonomy’s coverage of both cybersecurity threats and non-traditional content. Compared to traditional topic modeling, the LLM-guided approach provided superior contextual accuracy, offering a scalable framework for threat intelligence. Despite limitations, such as potential LLM biases and a focus on three forums, this study advances dark web analysis by refining the MISP taxonomy and revealing social dynamics alongside illicit activities.
This study presents an iterative methodological approach for the extension and evaluation of taxonomies applied to peer-to-peer (P2P) cryptocurrency forums in the dark web, a domain characterised by high semantic homogeneity and low thematic diversity. Drawing on a corpus of 23,642 posts, a rule-based deterministic pipeline was implemented to design an initial taxonomy focused on transactional intention, traded assets and payment mechanisms. Subsequently, non-classifiable cases were examined statistically using n-grams to detect empirical evidence of emerging subclasses, which led to the incorporation of new categories (primarily exchange-platform and forum) derived from the behaviour of the corpus itself. This extension eliminated ambiguous categories (other/unclear) and substantially improved classification coverage, showing that taxonomies in P2P domains should be understood as evolutionary artefacts rather than fixed structures.Under this definition, the present taxonomy evolves through a documented cycle of deterministic classification, ambiguity diagnosis, candidate extraction from residual cases, and constrained human-in-the-loop consolidation.Finally, co-occurrence analysis and lexical clustering revealed four differentiated functional communities, confirming that these forums do not operate as broad discursive spaces but as transactional infrastructures oriented towards operational execution. The results provide a traceable framework for explainable semantic classification in crypto environments, with direct implications for cyber intelligence and digital forensic analysis.
The Tor darkmarket ecosystem, a hidden segment of the internet hosting a range of illicit activities, remains a critical challenge for cybersecurity and law enforcement. This study employs network analysis to explore the structure, connectivity, and vulnerabilities of Tor hidden services, focusing on the interplay of topics, communication channels, and languages. Using a bipartite network framework, we analyzed 82,285 onion services and 57,071 identification forms (IDs) collected over a 20-week period. Our findings reveal hacking as the dominant topic (57,233 services), followed by finance-crypto (17,900 services), with email (43,298 IDs) and Telegram (11,218 IDs) serving as primary communication channels. Linguistically, Russian prevails in hacking (50,852 services), while English dominates other topics (29,762 services), with Portuguese activity notable in Q&A forums (781 services). Network metrics and visualizations highlight structural contrasts: hacking's expansive, collaborative structure (high diameter, long average path length) contrasts with finance-crypto's compact, centralized network (high density, low path length), reliant on just four IDs to link its services. High-degree nodes underscore vulnerabilities to targeted disruptions. The overall network's fragmentation (1848 components) alongside a large dominant component (76.72 %) suggests both resilience and exploitable interconnectedness. These insights provide a comprehensive understanding of the Tor darkmarket's organization, identifying key leverage points for intervention. By bridging gaps in topical, linguistic, and structural analyses, this study offers actionable strategies for law enforcement to investigate and mitigate illicit activities on the Dark Web, demonstrating the power of network science in addressing cybercrime.
Artificial intelligence language models are now embedded in social science workflows, yet the latent psychological and political patterns reflected in their outputs remain largely unexplored. This study examines 14 state-of-the-art large language models (LLMs) with 1000 prompt-based responses each (N=14,000) across three linked studies. Study 1 extracts the Big Five personality traits: On average, the outputs are characterized by high levels of agreeableness (M=6.54, SD=0.31) and conscientiousness (M=6.59, SD=0.35), very low neuroticism (M=1.57, SD=0.48), high openness (M=6.16, SD=0.53), and moderate extraversion (M=5.08, SD=0.83). Study 2 uses hierarchical regressions to predict authoritarianism and conspiracy mentality. Personality traits add on average only around 0.4% incremental R2, whereas political antecedents dominate. Study 3 measures ideological self-placement and sentiment toward Spanish leaders. Twelve models cluster at the partisan midpoint (M≈4), yet Mistral Medium leans toward the Spanish Socialist Party (PSOE, center-left) (M=3.32). Sentiment is neutral-positive for Pedro Sánchez (PSOE, center-left) and Yolanda Díaz (Sumar, far left), moderate for Alberto Núñez Feijóo (PP, center-right), and markedly negative for Santiago Abascal (Vox, far right) (sentiment=2.98; ideology=6.25). Overall, the outputs of LLMs converge in personality trait patterns but diverge sharply in how they translate into authoritarian and conspiratorial mentality. This study contributes the first exploratory analysis of personality traits and political orientations in AI language models and offers a replicable protocol for future comparative research.
Topic modeling is a critical tool for understanding the thematic structures of unstructured text data, particularly in specialized domains like the dark web. This study compares the effectiveness of large language models (LLMs) and traditional topic modeling techniques in analyzing dark web Q&A forums, characterized by short, informal, and context-specific posts. We evaluate two LLMs—GPT and Gemini—against traditional methods, TF-IDF (scikit-learn) and Latent Dirichlet Allocation (LDA, Gensim), in their ability to generate meaningful and coherent topics. Our findings suggest that LLMs consistently outperform traditional methods, capturing contextually relevant themes such as “scam,” “bitcoin,” and “hacking” with an average semantic similarity score of 0.724 between GPT and Gemini, compared to 0.477 with Gensim. Additionally, LLMs produced significantly lower Levenshtein distances, with an average of 13.49 between GPT and Gemini, compared to 18.22 with scikit-learn and 17.97 with Gensim, indicating greater alignment in topic extraction. In exploring the impact of topic granularity, we found that while single-word topics effectively capture the main themes of dark web discussions, two-word topics provide additional context and specificity, enhancing the interpretability of complex discussions. For instance, topics like “bitcoin wallet” and “hacking tools” illustrate how two-word labels can convey nuanced meanings that single-word topics may overlook. This qualitative analysis underscores the importance of selecting appropriate topic modeling techniques based on the characteristics of the text and the requirements of the analysis, highlighting their potential as valuable tools for analyzing complex and unstructured text data in practical applications for law enforcement and research.
This study investigates gender inequalities in academia by examining differences in representation, citations, and h-index between male and female highly cited researchers across disciplines and geographic regions. Using a unique dataset from Google Scholar, this study analyzes 21,509 highly cited authors across 191 fields and all continents. We examine gender disparities in citations, h-index, and representation while controlling for research productivity and career length to determine if female researchers experience different outcomes compared to their male counterparts. The findings reveal that women are significantly underrepresented among highly cited scholars globally (0.255 women per man) and receive fewer citations and have lower h-indexes than men in most regions and disciplines. However, after controlling for productivity and career length, female scholars are cited more than men in the pooled sample, Asia, Europe, and in two fields (natural sciences and exact sciences/physics). Despite this, women's h-index remains significantly lower than men's in all regions except Africa and South America, and in all fields except social sciences. This study highlights the persistence of gender inequalities in academic representation and long-term impact, as measured by the h-index. The results suggest that while citation rates for female researchers can match or exceed those of male scholars when productivity is controlled for, structural barriers continue to limit women's long-term recognition in academia. This research contributes to the understanding of gender disparities among top researchers, showing that while citation parity is possible, significant gender gaps remain in overall academic representation and long-term recognition through h-index measures.
This paper presents a protocol for using ChatGPT to perform content analysis. The protocol involves converting a codebook, outlining categories, descriptions, coding rules, and possible values, into a structured prompt that guides ChatGPT's analysis. The protocol was validated through analysis of 980 research articles to identify research approaches and data collection methods. ChatGPT achieved high performance in identifying data collection methods, but faced challenges with poorly defined or underrepresented categories, particularly in mixed methods research. Overall, while it scored well for quantitative (0.96) and qualitative (0.82) studies, it struggled with mixed methods (0.60), highlighting the need for clear methodological definitions.•The protocol enhances coding efficiency and demonstrates the feasibility of using AI for content analysis, potentially streamlining the coding process in research.•Challenges arose in categories that were not clearly defined (big data), underrepresented (ethnography), or hierarchically related (Interview & Discourse/Textual analysis).•Interrater metrics indicated a substantial level of agreement, reinforcing the potential of ChatGPT in content analysis while emphasizing the importance of clear methodological definitions.
The manufacturing sector's increasing reliance on Industry 4.0 technologies has made it a prime target for ransomware attacks, which can disrupt operations, cause financial losses, and compromise intellectual property. While prior studies have explored ransomware threats to industrial systems, few have leveraged dark web-disclosed data to understand the scale and nature of these attacks. This study analyzes 7,427 ransomware attack records disclosed on dark web onion services from April 2022 to March 2025, focusing on the manufacturing sector. The dataset, initially comprising 10,000 records, was cleaned by removing duplicates and records with missing NAICS codes, inferred using an AI-based approach. Findings reveal that manufacturing is vulnerable, with 1,620 attacks (21.81% of the total), tied with Professional Services as the most targeted sector. The United States accounted for 49.78% of manufacturing attacks, followed by Germany (7.10%), reflecting their significant manufacturing bases. A diverse set of 88 ransomware groups (78.57% of the total 112) targeted manufacturing, with LockBit responsible for 22.41% of attacks. These results underscore the urgent need for tailored cybersecurity strategies in manufacturing, including enhanced OT security and international collaboration to mitigate ransomware threats, particularly in high-risk regions like the U.S. and Germany.
The Dark Web hosts a variety of Q&A forums that facilitate discussions on topics ranging from technical guidance to illicit activities. This study investigates the linguistic and structural characteristics of three prominent Dark Web Q&A forum onion services—Onion1 (Deepweb Questions and Answers), Onion2 (Deep Answers), and Onion3 (Repostas Ocultas)—using a combination of topic modeling and quantitative text analysis. We employ Sklearn (TF-IDF) and Gensim (LDA) to extract and compare dominant topics, alongside metrics of lexical diversity, semantic diversity and syntactic complexity. Our findings reveal significant differences in topic diversity and linguistic patterns across the forums, with Onion1 exhibiting a focus on technical and financial discussions, Onion2 emphasizing community-driven and security-related topics, and Onion3 showing a narrower thematic scope influenced by its structured Q&A format and language-specific content. Statistical analyses, including Kruskal-Wallis tests and Dunn’s post-hoc comparisons, confirm these differences, highlighting the distinct communicative norms and user behaviors of each forum. The results underscore the importance of considering forum-specific characteristics when analyzing Dark Web communities and provide insights into the thematic and linguistic diversity of these platforms.
This study evaluates the zero-shot classification performance of eight commercial large language models (LLMs), GPT-4o, GPT-4o Mini, GPT-3.5 Turbo, Claude 3.5 Haiku, Gemini 2.0 Flash, DeepSeek Chat, DeepSeek Reasoner, and Grok, using the CoDA dataset (n = 10,000 Dark Web documents). Results show strong macro-F1 scores across models, led by DeepSeek Chat (0.870), Grok (0.868), and Gemini 2.0 Flash (0.861). Alignment with human annotations was high, with Cohen's Kappa above 0.840 for top models and Krippendorff's Alpha reaching 0.871. Inter-model consistency was highest between Claude 3.5 Haiku and GPT-4o (kappa = 0.911), followed by DeepSeek Chat and Grok (kappa = 0.909), and Claude 3.5 Haiku with Gemini 2.0 Flash (kappa = 0.907). These findings confirm that state-of-the-art LLMs can reliably classify illicit content under zero-shot conditions, though performance varies by model and category.
The Dark Web, a hidden segment of the internet, has become a hub for illicit activities, facilitated by various forms of digital identification (IDs) such as email addresses, Telegram accounts, and cryptocurrency wallets. This study conducts a comprehensive analysis of the Dark Web’s identification and communication patterns, focusing on the roles of different ID types and their associated activities. Using a dataset of Dark Web documents, we construct and analyze a bipartite network to model the relationships between IDs and web documents, employing graph–theoretical metrics such as degree centrality, closeness centrality, betweenness centrality, and k-core decomposition, while analyzing subnetworks formed by ID type. Our findings reveal that Telegram forms the backbone of the network, serving as the primary communication tool for hacking-related activities, particularly within Russian-speaking communities. In contrast, email plays a more decentralized role, facilitating finance–crypto and other activities but with a high level of fragmentation and English as the predominant language. XMR (Monero) wallets emerge as a key component in financial transactions, forming a cohesive subnetwork focused on cryptocurrency-related activities. The analysis also highlights the modular and hierarchical nature of the Dark Web, with distinct clusters for hacking, finance–crypto, and drugs–narcotics, often operating independently but with some cross-topic interactions. This study provides a foundation for understanding the Dark Web’s structure and dynamics, offering insights that can inform strategies for monitoring and mitigating its risks.
The gender classification from names is crucial for uncovering a myriad of gender-related research questions. Traditionally, this has been automatically computed by gender detection tools (GDTs), which now face new industry players in the form of conversational bots like ChatGPT. This paper statistically tests the stability and performance of ChatGPT 3.5 Turbo and ChatGPT 4o for gender detection. It also compares two of the most used GDTs (Namsor and Gender-API) with ChatGPT using a dataset of 5,779 records compiled from previous studies for the most challenging variant, which is the gender inference from full name without providing any additional information. Results statistically show that ChatGPT is very stable presenting low standard deviation and tight confidence intervals for the same input, while it presents small differences in performance when prompt changes. ChatGPT slightly outperforms the other tools with an overall accuracy over 96%, although the difference is around 3% with both GDTs. When the probability returned by GDTs is factored in, differences get narrower and comparable in terms of inter-coder reliability and error coded. ChatGPT stands out in the reduced number of non-classifications (0% in most tests), which in combination with the other metrics analyzed, results in a solid alternative for gender inference. This paper contributes to current literature on gender detection classification from names by testing the stability and performance of the most used state-of-the-art AI tool, suggesting that the generative language model of ChatGPT provides a robust alternative to traditional gender application programming interfaces (APIs), yet GDTs (especially Namsor) should be considered for research-oriented purposes.
Research is a global enterprise underpinned by the general belief that findings need to be true to be considered scientific. In the complex system of scientific validation, editorial boards (EBs) play a fundamental role in guiding journals’ review process, which has led many stakeholders of sciences to metaphorically picture them as the “gatekeepers of knowledge.” In an attempt to address the academic structure that governs sciences through editorial board interlocking (EBI, the cross-presence of EB members in different journals) and social network analysis, the aim of this study is threefold: first, to map the connection between fields of knowledge through EBI; second, to visualize and empirically test the distance between social and general sciences; and third, to uncover the institutional structure (i.e., universities) that governs these connections. Our findings, based on the dataset collected through the Open Editors initiative for the journals indexed in the JCR, revealed a substantial level of collaboration between all fields, as suggested by the connections between EBs. However, there is a statistically significant difference between the weight of the edges and the path lengths connecting the fields of natural sciences to the fields of social sciences (compared to the connections within), indicating the development of different research cultures and invisible colleges in these two research areas. The results also show that a central group of US institutions dominates most journal EBs, indirectly suggesting that US scientific norms and values still prevail in all fields of knowledge. Overall, our study suggests that scientific endeavor is highly networked through EBs.
Both computational social scientists and scientometric scholars alike, interested in gender-related research questions, need to classify the gender of observations. However, in most public and private databases, this information is typically unavailable, making it difficult to design studies aimed at understanding the role of gender in influencing citizens’ perceptions, attitudes, and behaviors. Against this backdrop, it is essential to design methodological procedures to infer the gender automatically and computationally from data already provided, thus facilitating the exploration and examination of gender-related research questions or hypotheses. Researchers can use automatic gender detection tools like Namsor or Gender-API, which are already on the market. However, recent developments in conversational bots offer a new, still relatively underexplored, alternative. This study offers a step-by-step research guide, with relevant examples and detailed clarifications, to automatically classify the gender from names through ChatGPT and two partially free gender detection tool (Namsor and Gender-API). In addition, the study provides methodological suggestions and recommendations on how to gather, interpret, and report results coming from both platforms. The study methodologically contributes to the scientometric literature by describing an easy-to-execute methodological procedure that enables the computational codification of gender from names. This procedure could be implemented by scholars without advanced computing skills.
High-performance computing (HPC) in data centers increases energy use and operational costs. Therefore, it is necessary to efficiently manage resources for the sustainability of and reduction in the carbon footprint. This research analyzes and optimizes ENEA HPC data centers, particularly the CRESCO6 cluster. The study starts by gathering and cleaning extensive datasets consisting of job schedules, environmental conditions, cooling systems, and sensors. Descriptive statistics accompanied with visualizations provide deep insight into collated data. Inferential statistics are then used to investigate relationships between various operational variables. Finally, machine learning models predict the average hot-aisle temperature based on cooling parameters, which can be used to determine optimal cooling settings. Furthermore, idle periods for computing nodes are analyzed to estimate wasted energy, as well as for evaluating the effect that idle node shutdown will have on the thermal characteristics of the data center under consideration. It closes with a discussion on how statistical and machine learning techniques can improve operations in a data center by focusing on important variables that determine consumption patterns.
The EU research project between industry and academia mu DevOps is a collaborative research project formed by an international network of organizations including industry and academia that aims to tackle current challenges of microservice development operations. An important case study considered in this project is the Cyber Ranges application a cyber security training and capability development exercises using microservices for the design, delivery, and management of simulation-based, experiences in cyber security developed by Silensec as one of the partners. This work describes the results of analyzing the scenario usage dataset of the Cyber Ranges training platform. This includes the matrix of starts for scenario/user and the attributes of scenarios. The aims are to produce recommendations of scenarios for users based on previous activity and to predict the success of scenarios as measured by the number of starts.
Microservice architectures are becoming increasingly important since they facilitate agile and modular production cycles to deliver applications using collections of loosely coupled and fine-grained services. The mu DevOps is a research project formed by an international network of organizations including industry and academia that aims to tackle current challenges of microservice development operations. This paper presents the mu DevOps project and the initial research carried out to evaluate the user experience of a microservice web application that delivers cybersecurity learning. Results point to critical elements of three main functionalities: library of scenarios, scenario information and entering scenario. Since there are several currently available solutions that offer similar services, the user experience of the microservice web app may play a critical role in determining which application will get a dominant role in the market.