This study investigates whether large language models (LLMs) can serve as reliable evaluators of machine translation (MT) quality for a low-resource, morphologically rich language-Slovak. Using $\mathbf{1 0 0}$ English-Slovak segments, we compare the fluency, adequacy, and usability ratings produced by seven human annotators and five LLM configurations (GPT-4o, GPT-4.1, and GPT-5 under multiple temperature settings). Statistical analyses reveal significant differences between human and LLM ratings across all criteria, but also show that advanced LLMs - particularly GPT-5 - exhibit strong internal consistency and form homogeneous groups with human annotators in fluency and usability assessment. Adequacy ratings demonstrate weaker alignment, reflecting both human disagreement and linguistic complexity. Overall, our results suggest that LLMs offer a scalable and reproducible complement to human MT evaluation, with near-human parity emerging in specific dimensions.
This study investigates the usability of machine translation (MT) outputs into Slovak, a low-resource and morphologically rich language, by examining the alignment between human evaluations and automated, reference-less quality estimation (QE) models. It aims to address the gap in MT evaluation that typically focuses on adequacy and fluency rather than practical usability. A QE model incorporating fluency, adequacy, and complexity features was compared against human judgments. The statistical analysis revealed a moderate, significant correlation, which remained stable even after excluding an outlier evaluator. Inter-annotator agreement was generally low to moderate, highlighting the variability in human assessments. Nevertheless, multiple comparisons indicated no significant differences between the QE model and five of the seven human evaluators, suggesting the model aligns closely with expert judgment. These findings underscore the challenges of human MT evaluation, particularly subjective variability inherent in assessing translation usability, while demonstrating the potential of QE models to support scalable, high-quality MT assessment. This study is the first to explore both human consistency and QE alignment for Slovak, offering valuable insights for improving MT evaluation in low-resource, morphologically complex languages.
The rapid expansion of cryptocurrency markets has coincided with the growing prominence of social media platforms as influential channels for shaping investor sentiment. Among these platforms, YouTube has become a medium for disseminating investment opinions and behavioral signals. This study investigates the extent to which YouTube-derived features—such as video influence scores, sentiment embedded in video titles, and user engagement indicators—can enhance the prediction of Bitcoin price fluctuations. A novel dataset is compiled, covering the period from January 2015 to September 2025. Sentiment is assessed using a combination of transformer-based language models, while influence metrics are computed through engagement statistics adjusted for temporal decay and relevance. These features are combined with historical Bitcoin price data and applied within an XGBoost forecasting framework. The empirical findings suggest that augmenting price-based models with YouTube-related sentiment and engagement features yields a notable improvement in directional forecasting accuracy, outperforming price-only benchmarks by approximately 4 %. Moreover, the study highlights that YouTube-derived behavioral signals offer predictive insights that are not fully captured by conventional indicators such as Google Trends or the Crypto Fear and Greed Index.
This multi-source dataset was compiled to support research on anomaly-based leak detection in urban water distribution networks (WDNs). It contains one year of hourly data collected from a Slovak water utility, combining supervisory control and data acquisition (SCADA) measurements (flow and pressure), energy consumption variables (kilowatts, kW; kilovolt-ampere reactive, kvar), and environmental indicators such as groundwater level, temperature, and humidity. All features were transformed into standardized anomaly scores on a 0-100 scale using Elastic Machine Learning and Isolation Forest methods. Confirmed leak records from the utility's operational information system were mapped to binary labels using a ± 7-day temporal window. Feature selection resulted in 18 variables retained based on their statistical association with leak labels using the Goodman-Kruskal gamma coefficient. The dataset can be used for benchmarking anomaly detection and prediction models, evaluating lead-time sensitivity, and developing data-driven early-warning systems. It is also suitable for studies on spatially segmented WDN analysis and is publicly available to support reproducible research.
This article focuses on a comprehensive evaluation of three available tools—online services (ChatGPT, Google Cloud NLP, and NLP Cloud)—for determining sentiment from text. Currently, there is a strong emphasis on understanding not only the content of the text but also the manner (style and method) in which it is written, especially in interpersonal communication. We input 20 texts with varying sentiment tones into these services using an interface. Additionally, ten psychologists examined the texts to clearly determine, from a professional and human perspective, the feeling each text evokes in a reader. From the results, we observe insignificant differences in the scores among the tools investigated. The highest score was achieved by the Google Cloud tool (0.14), while the lowest scores were achieved by the NLP Cloud (0.10) and ChatGPT (0.09) tools. The highest score of variability was identified for NLP Cloud (0.97). Statistical results simultaneously imply that there is no statistically significant difference in the sentiment scores among individual psychologists for the positive sentiment level, according to Google Cloud results. The evaluations of individual psychologists collectively form one homogeneous group in terms of the sentiment score for the positive level, according to Google Cloud results. While repeated-measures tests were non-significant and confidence intervals overlapped across tools in our corpus, these findings indicate within-sample similarity rather than general interchangeability. We therefore refrain from interchangeability claims and interpret the results as no detected differences in this dataset, conditional on the texts, domains, and scoring scale used.
The paper builds research examining the implementation of digital didactic tools in pedagogical practice and focuses on the pilot verification of the research instrument used in this study. The questionnaire was administered to identify current needs of educational practice related to the development of teachers’ professional competencies within the Slovak regional school system, particularly their digital literacy and didactic-technological competencies affecting the quality and effectiveness of instruction. It was designed to provide valid data on the level of digital competencies of primary and secondary school teachers, especially in relation to their use of selected digital tools and technologies in teaching. The pilot verification aimed at assessing the instrument’s psychometric properties, with emphasis on reliability. The paper presents the results of reliability and item analyses. Internal consistency was examined using reliability coefficients, and potentially problematic items were identified through item analysis. The findings indicate satisfactory internal consistency of the analysed scales and confirm the instrument’s suitability for the main research phase, while suggesting minor refinements of selected items.
This paper presents a language-specific adaptation for automatic identification of machine translation (MT) errors using a comprehensive set of open-source evaluation metrics. The approach focuses on the English–Slovak translation direction, addressing challenges posed by Slovak’s highly inflectional and low-resource nature. Predictive models were developed for five key error categories (Predication, Modal and communication sentence framework, Syntactic-semantic correlativeness, Compound/complex sentences, and Lexical semantics) by employing forward stepwise regression and validated through bootstrapping techniques. The models estimate the probability of error occurrence in MT segments, demonstrating stable and comparable performance across training and test datasets, as measured by Somers’ D. While human expert evaluation remains essential for verifying flagged segments, the proposed approach significantly reduces evaluator workload by prioritizing likely error-containing segments. This methodology offers a scalable and adaptable framework for MT quality assessment across languages and text styles, with potential to improve automated translation evaluation and post-editing processes.
Financial statement fraud undermines market integrity and incurs substantial costs for investors, regulators, and companies. Text-based detection methods have emerged as useful complements to traditional financial indicators, but many fail to incorporate domain-specific topics or sentiment cues, often missing subtle changes in deceptive communication. To overcome this problem, this study proposes a topic-driven financial sentiment analysis (TDFSA) model that detects corporate fraud by analyzing linguistic patterns in the Management Discussion & Analysis (MD&A) sections of annual reports. Our approach captures contextual sentiment within financially relevant topics using FinBERT embeddings. To evaluate these signals in fraud detection, we integrate the TDFSA outputs into a broader cost-sensitive evaluation framework. This framework combines text-based indicators with financial ratios to balance the need to avoid false alarms with the high cost of undetected fraud. Using data from U.S. firms flagged in SEC Accounting and Auditing Enforcement Releases from 2014 to 2024 and matched non-fraud peers, we examine trends in financial ratios, textual complexity, and sentiment dynamics in the three years preceding fraud events. The results show that models leveraging TDFSA achieve higher detection accuracy and lower cost than dictionary-based sentiment, generic topic models, and deep learning baselines.
Introduction: Pillar 3 of the Basel framework aims to strengthen market discipline by requiring banks to disclose detailed information on risk exposures, capital adequacy, and liquidity. Existing research has primarily assessed stakeholder engagement with these disclosures through analyses of web server log files from commercial bank websites. While effective, such approaches are constrained by limited data accessibility, substantial computational demands, and restricted geographical scalability This study proposes a methodology for assessing stakeholder interest in Pillar 3 disclosures by combining multilingual NLP-based keyword extraction with search behaviour analytics based exclusively on publicly available data. Methods: The proposed methodology integrates natural language processing techniques for extracting representative keywords from Pillar 3–related disclosure documents with Google Trends data capturing search-based indicators of stakeholder interest. Keywords were extracted from multiple disclosure subcategories and translated into the languages of the Visegrad Four countries (Czech Republic, Hungary, Poland, and Slovakia). Search interest was analysed on a quarterly basis using repeated-measures analysis of variance, with quarter as the within-group factor and disclosure subcategory and country/language as between-group factors. Multiple weighting schemes - keyword frequency, keyword specificity, and keyword density - were applied to evaluate the robustness and discriminatory power of the results. Results: The results reveal statistically significant variation in stakeholder interest across quarters, disclosure subcategories, and countries. Search interest consistently peaks in the first and fourth quarters, reflecting financial reporting cycles Annual reports and covered bonds attract higher and more stable levels of interest compared to other disclosure categories. Significant cross-country differences are observed, with Poland exhibiting the highest overall level of search interest among the Visegrad Four. Conclusion: The study confirms that publicly available search data can serve as a reliable and scalable proxy for assessing stakeholder engagement with Pillar 3 disclosures. The proposed methodology offers a practical alternative to log file-based analyses, enabling regular, region-wide monitoring of disclosure effectiveness without reliance on proprietary data sources. The findings have direct implications for regulators and financial institutions seeking to evaluate and enhance the effectiveness of Pillar 3 disclosures in supporting market discipline and transparency.
Annual reports are important in globalised economies as they are serving as tools for communication between banks and their stakeholders, including shareholders/investors, analysts, regulators/government agencies, other financial institutions, employees, and business partners. This study utilises sentiment analysis to examine shifts in sentiment within Slovak annual reports across key periods: pre-financial crisis, financial crisis, post-financial crisis, post-revision of the third pillar, COVID-19 pandemic, and the energy crisis/war in Ukraine. Results reveal notable patterns, with the global financial crisis being the most turbulent period, marked by a higher ratio of negative sentences. In contrast, the Post-Pillar 3 revision period exhibits stability, featuring the highest ratio of neutral sentences. This research contributes valuable insights into corporate sentiment dynamics during challenging times, shedding light on how companies navigate uncertainty and present narratives in annual reports, influencing stakeholder perceptions and potentially signalling future corporate decisions and strategies.
Machine Translation (MT) evaluation plays a crucial role in advancing systems translating into morphologically rich, low-resource languages such as Slovak. Existing automatic evaluation methods typically offer a single quality score, lacking insight into specific error types. A novel linguistically informed methodology that predicts the probability of MT error categories by integrating manual annotation with automatic evaluation metrics is proposed. The method builds on a modified MQM framework adapted for Slovak and employs a dataset of English-to-Slovak translations, combining outputs from statistical and neural MT systems with human reference translations. Manual annotations identified five linguistically motivated error categories. Reliability of 68 automatic metrics was assessed using Cronbach’s alpha, correlation coefficients, coefficient of determination (R²), and entropy. Bootstrapped logistic regression models were then developed to predict error occurrence probabilities. The proposed methodology improves the explainability and reliability of automatic MT evaluation by bridging the gap between holistic scoring and detailed error categorization. It significantly reduces the human effort required for quality assessment while maintaining a high degree of linguistic relevance, particularly for complex target languages like Slovak. • Predicts probabilities of specific MT error categories • Integrates linguistic expertise with statistical reliability analysis • Reduces human effort in MT evaluation while preserving linguistic precision
One of outstanding issues of Pillar 3 information disclosure is to increase the relatively low interest of stakeholders of commercial banks operating in the CEE region in Pillar 3 information. The aim of this paper is to investigate the relations between the sentiment expressed in different sections of annual reports and the performance and risk indicators, disclosed as part of the Pillar 3 disclosures, of a commercial bank operating in CEE country. We examined risk and performance indicators in the context of textual disclosures of six sections of the annual reports. We found out that sections the supervisory board and board of directors presented the weakest dependence followed by external environment and results. The strongest dependence has been found in sections assessment and information on expected economic and financial situation in upcoming year.
The incidence of neurodegenerative diseases affecting the brain and its cognitive functions, including language and speech, is increasing in society. These diseases impact the manner and quality of speech and can be detected through non-invasive methods. Understanding language involves analyzing internal linguistic features such as text readability and complexity. Language complexity is a significant measure of an individual's linguistic development, representing an independent dimension of utterance (whether written or spoken) and manifesting across all linguistic levels (phonological, morphological, syntactic, and semantic). The aim of this research is to identify linguistic features – measures of text complexity – that may serve as predictors for a knowledge-based model to detect neurodegenerative diseases such as Alzheimer's Disease (AD), Mild Cognitive Impairment (MCI), and Parkinson's Disease (PD) in the context of the inflectional Slovak language. The results indicate that lexical measures of language complexity that are independent of text length are unsuitable for predicting neurodegenerative diseases such as AD/MCI or PD. However, they can be useful in distinguishing between AD/MCI and PD. The rate of action in describing a situational picture is a strong predictor for distinguishing AD/MCI but not PD. The sequence of two verbs is a strong predictor for diagnosing both AD/MCI and PD, but does not distinguish between these diseases. Last but not least, vocabulary range and diversity influence not only the diagnosis of neurodegenerative diseases, but also help differentiate between AD and PD.
This research explores the impact of document complexity and readability on user preferences for disclosure information on commercial bank web portals, with a focus on Pillar 3 disclosures. The study investigates the usage patterns of disclosed information by stakeholders. Through an analysis of web portal access variables and document complexity and readability measures, the study identifies correlations between text complexity/readability and user preferences. The experiment used traditional readability measures like the Gunning Fog Index and Flesch Reading Ease but also focused on less-used metrics such as LIX, various text characteristics, and part-of-speech metrics. The findings reveal that while stakeholders show interest in financial documents despite their complexity, preferences in using these documents vary based on document type and readability. Notably, documents with higher complexity tend to attract more attention, suggesting potential challenges in accessibility and comprehension. The paper highlights the importance of presenting disclosure information in a manner that enhances readability and accessibility to better engage stakeholders. Additionally, the research introduces user preference indicators aligned with web portal access metrics, providing insights into stakeholders’ behavior. The research followed up on previous findings of extremely low interest in Pillar 3 information from commercial bank stakeholders. The results showed that stakeholders are more interested in less demanding and more readable texts such as annual reports, and as Pillar 3 documents are more complex and harder to read, they are less interested in them. The research concludes that stakeholders’ increasing interest in this information requires finding approaches to presenting it in less complicated and more readable ways.
This study explores the communication patterns of Slovak banks with stakeholders through mandatory disclosures mandated by Basel III's Pillar 3 framework and annual reports in 2007-2022. Our primary objective is to identify key topics communicated by banks and analysing the sentiment of this communication during turbulent periods (i.e., alternating periods of stability and crisis) in 2007-2022. Textual data was collected from Pillar 3 disclosures, annual reports, and additional regulatory reports. A hybrid model was developed to extract the most important keywords from each collected document chapter. This hybrid model (model combining multiple approaches) combines elements of statistical approaches to keyword extraction, (keyword frequency dictionary), linguistic approaches (pair-of-speech tagging in order to select noun-phrases), and machine-learning based approaches (BERT) to extract meaningful keywords. Subsequently, a sentiment analysis was performed on the extracted keywords using a Loughran-McDonald lexicon (list of words labelled with sentiment) specially designed for financial texts. Based on the adjusted univariate results, we can reject the global null hypothesis of independence of the sentiment category of keywords from time for negative sentiment at p = 0.0000 for positive sentiment at p = 0.0005, and for neutral sentiment at p = 0.0000 significant level. The multilevel comparison revealed that negative sentiment was most frequent during the global financial crisis and the COVID-19 pandemic, likely impacting stakeholder confidence and trust. Conversely, positive sentiment dominated during periods of financial stability, potentially enhancing stakeholder satisfaction and investment decisions. This research points out that the sentiment of the selected commercial bank documents changes depending on the years. A commercial bank can use this knowledge and include sentiment information as predictors when modelling financial distress. For bank management of selected commercial bank the examined documents are an important communication tool, the wording of which can have a significant impact on stakeholder behaviour towards the bank, their styling is very important.
Water Distribution Systems (WDSs) are crucial for urban infrastructure but struggle with challenges such as significant water losses, with up to 40
The 2008–2009 financial crisis exposed significant flaws in the Basel II regulatory framework, which was designed to ensure the stability of banking institutions during crises. This economic turmoil led to a loss of confidence in financial systems, liquidity shortages, and a deepening debt crisis across Europe. In response, regulatory standards were reformed, resulting in the development of the Basel III framework, which built upon three pillars. Pillar 1 introduced better capital forms, higher capital ratios, and capital reserves. Pillar 2 increased risk management requirements and emphasized the supervisory authorities’ responsibilities. Pillar 3 enhanced the scope and detail of published information, aiming to improve market discipline and stakeholder understanding of banks’ risks and capital positions. Our study analyzes how crisis periods affected interest in Pillar 3 disclosures, examining changes in interest from 2016 to 2022 using data from a specific bank’s log files. The results, performed through association analysis, indicate a significant decrease in interest in mandatory disclosures during the periods of the COVID-19 pandemic as well as the energy crisis and the war in Ukraine.
BACKGROUND:Dementia, particularly Alzheimer's disease (AD), affects language, especially lexical-semantic processing. Discourse analysis using NLP methods can aid early detection, but research in inflectional languages like Slovak is limited. METHODS:Speech samples from 216 Slovak-speaking participants (64 AD, 44 MCI, 108 HC) were collected using a picture description task and analyzed for lexical complexity using 15 NLP-based measures. RESULTS:Several lexical complexity measures, including GTTR, UBER, SICHEL, MTLD, HDD and others, significantly differentiated AD or MCI from healthy controls. Some measures (UBER, YULEI, HONORE) also distinguished between AD and MCI. CONCLUSION:Lexical complexity metrics can serve as non-invasive linguistic indicators of neurodegenerative diseases, demonstrating diagnostic relevance for early detection of AD and MCI in Slovak. Highlights:Lexical complexity metrics effectively differentiate between healthy controls, MCI, and AD in Slovak speakers.Measures such as GTTR, UBER, and HONOR exhibit strong diagnostic potential for neurodegenerative diseases.Education significantly influences linguistic deficits, with higher education correlating to reduced cognitive decline.Findings underscore the importance of studying minority languages for advancing AD and MCI diagnostics.
This paper proposes a novel multimodal deep learning framework for stock return prediction that integrates heterogeneous data sources: technical indicators, market investor sentiment indices, and textual sentiment extracted from earnings conference call transcripts. The proposed model employs a hybrid architecture combining transformer encoder for the technical modality and neural networks for market and textual modalities. A modality-level attention mechanism is used in a late fusion setup to dynamically weight the contributions of each modality. We evaluate our model on a large-scale dataset comprising 24,821 samples from 497 S&P 500 companies over the period 2010–2022. The results show that our model outperforms traditional models (LSTM, BiL-STM, CNN-LSTM) and alternative fusion strategies, achieving a directional accuracy of 59.94% on the test set. Attention weight analysis confirms that all three modalities contribute meaningfully to prediction performance. These results demonstrate the overall effectiveness of the proposed framework in accurately predicting abnormal stock returns in a multimodal setting.
This data article releases geospatial predictor files for a section of the Gidra River (western Slovakia). Based on the last cycle of the Preliminary Flood Risk Assessment from 2024 in Slovakia, the past fluvial floods and, especially, the flash flood from 7 June 2011 affected the studied river section and the Píla municipality significantly. These facts resulted in including the studied Gidra River section to critical river sections for the occurrence of fluvial floods. The collection comprises seven single-band rasters on a common 1 m grid: slope, Topographic Wetness Index, Stream Power Index, Height Above the Nearest Drainage, Euclidean distance from river, surface roughness, and Normalized Difference Vegetation Index. Source inputs are the LiDAR digital elevation model (DMR 5.0, resolution: 1 m) from 2017 and aerial orthophotos (resolution: 0.15 m) from 2023. All layers are georeferenced to a single CRS, co-registered to identical extent and transform, and provided as GeoTIFFs with documented units, no data, and datatypes. Processing relied on ArcGIS 10.2.2 software and Python for alignment and quality checks. Additional tables supply per-file inventory and descriptive statistics (min/mean/max/std) to enable automated validation and integration into geographic information system (GIS) and modelling workflows. The dataset is designed for reuse in flood-related and terrain-vegetation analyses, including feature engineering, benchmarking, and training of machine-learning (ML) models that require uniform, high-resolution predictors.