Court judgments are central to legal practice and jurisprudence, yet discourse analysis of Hong Kong judgments has received limited attention, owing largely to the absence of expert-annotated corpora. We introduce the Hong Kong Judgment Discourse Dataset (HKJudge), the first sentence-level expert-annotated legal discourse corpus. HKJudge includes criminal judgments across all five levels of HK's court hierarchy, comprising ∼290k sentences and ∼6.5 million tokens, fully annotated by legal linguistics experts. We design a two-tier discourse schema that captures what facts a court finds, how it reasons, and what it rules. At the sentence level, each sentence is assigned one of 26 rhetorical roles. At the span level, sentences are further annotated with three sentencing elements (charge, imprisonment term, fine). Ten legal linguistics annotators produced the annotations with an inter-annotator agreement of κ= 0.8. We formulate two tasks on HKJudge, termed rhetorical role classification and legal element extraction, and provide the first benchmark evaluation of four BERT-based models, two open-source LLMs under zero-shot and fine-tuning settings, and four commercial LLMs on both tasks. Our work demonstrates the value of sentence-level discourse annotation for modeling the structure of HK judgments and provides a rich data foundation for future work on legal judgment prediction. The HKJudge dataset and code are available at https://github.com/xuanxixi/HKJudge.
This study investigates how wrap-up effects—sentence-final words incur heavier processing loads than sentence-internal words—manifest in natural reading of unspaced, logographic Chinese scripts. We leveraged a large-scale naturalistic reading corpus (subjects: 98; word tokens: more than 1 million by 300 individual sentences and seven passages), employed multifactorial analyses by (generalized) linear mixed-effects models, and compared four eye-movement differences (i.e., gaze duration, total reading time, skipping probability and regression-in probability) between sentence-internal and sentence-final words. The results demonstrated robust reversal of traditional wrap-up effects: sentence-final words required significantly less processing effort than sentence-internal words, such as shorter durations, higher skipping rates and lower regression-in probabilities. Further, this reversal was modulated by reading scenarios (sentence vs. passage), boundary salience (period- vs. comma-bounded), and wrap-up positions (pre-critical, critical, and spill-over). Notably, sentence-final words were processed more rapidly when associated with characteristics such as fewer stroke counts, shorter length, higher frequency, function-word status, or progress further into page at the late processing stage. Challenging classic models that attribute inflated time at clause/sentence boundaries to semantic integration, we postulate punctuation’s dual role in unspaced language processing: visual cues for word segmentation/recognition and semantic cues for integration jointly optimize language comprehension.
Hong Kong case law translation presents significant challenges: manual methods suffer from high costs and inconsistent quality, while both traditional machine translation and approaches relying solely on Large Language Models (LLMs) often fail to ensure legal terminology accuracy, culturally embedded nuances, and strict linguistic structures. To overcome these limitations, this study proposes TransLaw, a multi-agent framework that decomposes translation into word-level expression, sentence-level translation, and multidimensional review, integrating a specialized Hong Kong legal glossary database, Retrieval-Augmented Generation (RAG), and iterative feedback. Experiments on our newly constructed HKCFA Judgment 97-22 dataset, benchmarking 13 open-source and commercial LLMs, demonstrate that TransLaw significantly outperforms single-agent baselines across all evaluated models. Human evaluation confirms the framework's effectiveness in terms of legal semantic accuracy, structural coherence, and stylistic fidelity, while noting that it still trails human experts in contextualizing complex terminology and stylistic naturalness.
This paper addresses the challenges translating case law under Hong Kong's bilingual legal system. It highlights the initial success of translating all written statutes into Chinese before the 1997 handover, a task mandated by the Basic Law. The effort involved significant collaboration among legal, linguistic, and translation experts, resulting in a comprehensive and culturally appropriate bilingual legal system. However, translating case law remains a significant challenge due to the sheer volume and continuous growth of judicial decisions. The paper critiques the governments and judiciarys sporadic and uncoordinated efforts to translate case law, contrasting it with the thorough approach previously taken for statute translation. Although the government acknowledges the importance of legal bilingualism, it lacks a sustainable strategy for translating case law. The Judiciarys position that translating all judgments is unnecessary, unrealistic, and not cost-effectiveis analyzed and critiqued for its impact on legal transparency and public trust. A proposed solution involves leveraging machine translation technology through a human-machine interactive translation platform, which undergoes two major transitions. Initially based on a neural model, the platform transitions to using a large language model for improved translation accuracy. Furthermore, it evolves from a single-agent system to a multi-agent system, incorporating Translator, Annotator, and Proofreader agents. This multi-agent approach, supported by a grant, aims to facilitate efficient, high-quality translation of judicial judgments by integrating advanced artificial intelligence and continuous feedback mechanisms, thus better meeting the needs of a bilingual legal system.
Materials discovery and design aim to find compositions and structures with desirable properties over highly complex and diverse physical spaces. Traditional solutions, such as high-throughput simulations or machine learning, often rely on complex descriptors, which hinder generalizability and transferability across different material systems. Moreover, these descriptors may inadequately represent macro-scale material properties, which are influenced by structural imperfections and compositional variations in real-world samples, thus limiting their practical applicability. To address these challenges, we propose DARWIN 1.5, the largest 1 open-source large language model tailored for materials science. By utilizing natural language as input, DARWIN eliminates the need for task-specific descriptors and facilitates the integration of human knowledge representation with computational models, enabling a more flexible and unified approach to material property prediction and discovery. Our approach integrates over 6M materials science papers and 21 experimental datasets with information of 49,256 materials, allowing for efficient cross-task knowledge transfer and improved generalization. Through systematic exploration, we show how domain-specific know-how can be effectively integrated into language models while harnessing the inherent syn-ergies between tasks to enhance predictive performance across diverse material science applications. The enhanced model achieves up to 59.1% improvement in prediction accuracy over the base LLaMA-7B model architecture and outper-forms state-of-the-art machine learning approaches across eight materials design tasks. These results highlight the potential of LLMs as a foundation for developing versatile and scalable models in materials science.
The enduring comparison between Li Bai ((sic)(sic), 701-762) and Du Fu ((sic)(sic), 712-770), two towering poets in Chinese literary history, has traditionally centred on semantic and thematic aspects, often overlooking the phonetic dimension or narrowly focusing on rhythm and rhyme while neglecting other sound features. This study applies sound symbolism principles to explore the emotional and perceptual effects of sounds in their poetry, focusing on a quantitative phonetic comparison of their works. Utilising character sound information extracted from the rhyme book Guangyun ((sic)(sic) extended rhymes), their poems were transformed into sound vectors for authorship attribution by machine learning models. Several of these models achieved average F1 scores above 80%, effectively demonstrating the capability of sounds to distinguish between the two poets. Through statistical analyses, their preferred phonetic features were located and the sound-sentiment relationships in their poetry were identified, revealing that Li Bai favoured the more positive sound features while Du Fu showed an inclination towards the more negative ones. Moreover, the distinct sound preferences sculpted Li's poetry with a vibrant, melodious, and unbound quality, while lending a subtle, sombre, and constrained undertone to Du's verses. Focusing on the phonetic aspect, this study takes a novel digital humanities approach to the traditional Li-Du comparison topic, offering fresh insights into their differences and underscoring the impact of sounds on the aesthetic experience of poetry.
Contrary to the widespread notion that linguistic signs are arbitrary, researchers have consistently demonstrated the existence of sound symbolism in language, providing evidence for non-arbitrariness in sound-meaning associations. However, much evidence of this kind is based on a limited subset of vocabulary and falls short of systematically demonstrating the pervasive nature of sound symbolism and, especially, its central, rather than marginal, role in language. Furthermore, a historical perspective is lacking to determine whether sound symbolism is merely a feature of archaic languages or has remained a significant element throughout the evolution of languages. This research pioneers a diachronic analysis of sound symbolism in Chinese using historical rhyme books to trace its presence on the vocabulary scale. Employing natural language processing techniques along with statistical methods, it investigates whether phonologically related Chinese characters, as documented in rhyme books, also demonstrate semantic congruence, which would suggest that the phonological aspects of characters are inherently meaningful and hence indicate a systematic, rather than random or purely arbitrary relationship between sounds and meanings. Statistically significant results from our analysis of all four analyzed rhyme books confirm the robustness of sound symbolism over a large span of the Chinese language continuum, and a granular analysis of a representative one of them further reveals that sound symbolism is manifest across various levels of phonological organization, including initials, finals, etc. This study initiates an innovative combination of traditional materials with novel techniques to enrich and expand existing knowledge about sound symbolism, providing both methodological advancement and empirical insights.
Exploring the predictive capabilities of language models in material science is an ongoing interest. This study investigates the application of language model embeddings to enhance material property prediction in materials science. By evaluating various contextual embedding methods and pre-trained models, including Bidirectional Encoder Representations from Transformers (BERT) and Generative Pre-trained Transformers (GPT), we demonstrate that domain-specific models, particularly MatBERT significantly outperform general-purpose models in extracting implicit knowledge from compound names and material properties. Our findings reveal that information-dense embeddings from the third layer of MatBERT, combined with a context-averaging approach, offer the most effective method for capturing material-property relationships from the scientific literature. We also identify a crucial "tokenizer effect," highlighting the importance of specialized text processing techniques that preserve complete compound names while maintaining consistent token counts. These insights underscore the value of domain-specific training and tokenization in materials science applications and offer a promising pathway for accelerating the discovery and development of new materials through AI-driven approaches.
We introduce SciQAG, a novel framework for automatically generating high-quality science question-answer pairs from a large corpus of scientific literature based on large language models (LLMs). SciQAG consists of a QA generator and a QA evaluator, which work together to extract diverse and research-level questions and answers from scientific papers. Utilizing this framework, we construct a large-scale, high-quality, open-ended science QA dataset containing 188,042 QA pairs extracted from 22,743 scientific papers across 24 scientific domains. We also introduce SciQAG-24D, a new benchmark task designed to evaluate the science question-answering ability of LLMs. Extensive experiments demonstrate that fine-tuning LLMs on the SciQAG dataset significantly improves their performance on both open-ended question answering and scientific tasks. To foster research and collaboration, we make the datasets, models, and evaluation codes publicly available, contributing to the advancement of science question answering and developing more interpretable and reasoning-capable AI systems.
Over the past five decades, onomastics has seen remarkable growth with fruitful publications and interdisciplinary collaborations. Despite the abundance of literature, a panoramic view of contribution networks and the evolutionary trajectory of this field has been lacking. To address this issue, this study presents a statistical assessment complemented by visualization clustering, rendering data from 768 journal articles and 28,357 references, to unfold impactful journals, influential scholars, foundational knowledge, and evolving frontiers. The outcomes of this research showcase the distribution of subtopics within each name category, depicting noteworthy contributors, focal trends, and cutting-edge subjects in the area. New themes that illuminate orientations include online naming, multi-identity construction, language processing, corpus-assisted approaches, and neural-cognitive experiments. Further data-driven exploration of name-related themes is foreseen to yield valuable insights. Through this comprehensive assessment, this study elucidates the role of names as manifestations of human identity, social emotions, aesthetic ingenuity, and strategic communicative paradigms. The findings are poised to facilitate the discernment of human quality, societal stratification, interpretative nuances, and relationships underlying social issues. Additionally, this research exemplifies the efficacy of bibliometric analysis and proposes strategies to mitigate potential constraints, disclosing how quantitative data from onomastics can be applied in the digital era and beyond.
Materials scientists usually collect experimental data to summarize experiences and predict improved materials. However, a crucial issue is how to proficiently utilize unstructured data to update existing structured data, particularly in applied disciplines. This study introduces a new natural language processing (NLP) task called structured information inference (SII) to address this problem. We propose an end-to-end approach to summarize and organize the multi-layered device-level information from the literature into structured data. After comparing different methods, we fine-tuned LLaMA with an F1 score of 87.14% to update an existing perovskite solar cell dataset with articles published since its release, allowing its direct use in subsequent data analysis. Using structured information, we developed regression tasks to predict the electrical performance of solar cells. Our results demonstrate comparable performance to traditional machine-learning methods without feature selection and highlight the potential of large language models for scientific knowledge acquisition and material development.
In the studies of classical Chinese poetry, the comparison between Li Bai and Du Fu is an everlasting topic, yielding many qualitative interpretations, among which a widely known but disputable one is Li's positivity versus Du's negativity. With the development of digital means, distant reading has become possible, and the sentiment issue can be further explored in quantitative ways. This research conducts a corpus-based sentiment comparison of Li and Du with a self-constructed sentiment dictionary. The Complete Collection of Tang Poems is used as a representative of Tang poets, and sentiment comparisons are made at the levels of poems, verses, and characters, as well as key characters extracted with the log-likelihood measure. Analyses show that (1) among Tang poets, Du is more negative at all of the above textual levels, while Li is only more positive at the key character level, proving the importance of key characters in readers' perception of sentiment; (2) Li and Du both stand out among Tang poets with a negative depiction of the dark reality and a positive expression of grand ideals; and (3) Li's positivity is largely embodied in his depictions of color, light, and temperature, while Du's negativity is closely related to his psychological description. To conclude, this research has not only determined the sentiment difference between Li and Du but also located its sources in texts with a novel key character-based sentiment analysis approach.Keywords: Li Bai, Du Fu, sentiment analysis, corpus, keyness
Natural Language Processing (NLP) is widely used to supply summarization ability from long context to structured information. However, extracting structured knowledge from scientific text by NLP models remains a challenge because of its domain-specific nature to complex data preprocessing and the granularity of multi-layered device-level information. To address this, we introduce ByteScience, a non-profit cloud-based auto fine-tuned Large Language Model (LLM) platform, which is designed to extract structured scientific data and synthesize new scientific knowledge from vast scientific corpora. The platform capitalizes on DARWIN, an open-source, fine-tuned LLM dedicated to natural science. The platform was built on Amazon Web Services (AWS) and provides an automated, user-friendly workflow for custom model development and data extraction. The platform achieves remarkable accuracy with only a small amount of well-annotated articles. This innovative tool streamlines the transition from the science literature to structured knowledge and data and benefits the advancements in natural informatics. Demo Video
Materials discovery and design aim to find compositions and structures with desirable properties over highly complex and diverse physical spaces. Traditional solutions, such as high-throughput simulations or machine learning, often rely on complex descriptors, which hinder generalizability and transferability across different material systems. Moreover, These descriptors may inadequately represent macro-scale material properties, which are influenced by structural imperfections and compositional variations in real-world samples, thus limiting their practical applicability. To address these challenges, we propose DARWIN 1.5, the largest open-source large language model tailored for materials science. By leveraging natural language as input, DARWIN eliminates the need for task-specific descriptors and enables a flexible, unified approach to material property prediction and discovery. Our approach integrates 6M material domain papers and 21 experimental datasets from 49,256 materials across modalities while enabling cross-task knowledge transfer. The enhanced model achieves up to 59.1 accuracy over the base LLaMA-7B architecture and outperforms SOTA machine learning approaches across 8 materials design tasks. These results establish LLMs as a promising foundation for developing versatile and scalable models in materials science.
Emerging tools bring forth fresh approaches to work, and the field of natural science is no different. In natural science, traditional manual, serial, and labour-intensive work is being augmented by automated, parallel, and iterative processes driven by artificial intelligence-based experimental automation and more. To add new capabilities in natural science, enabling the acceleration and enrichment of automation of the discovery process, we present DARWIN, a series of tailored LLMs for natural science, mainly in physics, chemistry, and material science. This series relies on open-source LLM, incorporating structured and unstructured scientific knowledge from public datasets and literature. We fine-tuned the models using over 60,000 instruction data points, emphasizing factual correctness. During the fine-tuning, we introduce the Scientific Instruction Generation (SIG) model, automating instruction generation from scientific texts. This eliminates the need for manual extraction or domain-specific knowledge graphs and efficiently injects scientific knowledge into the model. We also explore multi-task training strategies, revealing interconnections between scientific tasks. DARWIN series not only achieves state-of-the-art results on various scientific tasks but also diminishes reliance on closed-source AI models. Our research showcases the ability of LLM in the scientific domain, with the overarching goal of fostering prosperity within the broader AI for science community.
Recent years have witnessed a mushrooming of reading corpora that have been built by means of eye tracking. This article showcases the Hong Kong Corpus of Chinese Sentence and Passage Reading (HKC for brevity), featured by a natural reading of logographic scripts and unspaced words. It releases 28 eye-movement measures of 98 native speakers reading simplified Chinese in two scenarios: 300 one-line single sentences and 7 multiline passages of 5,250 and 4,967 word tokens, respectively. To verify its validity and reusability, we carried out (generalised) linear mixed-effects modelling on the capacity of visual complexity, word frequency, and reading scenario to predict eye-movement measures. The outcomes manifest significant impacts of these typical (sub)lexical factors on eye movements, replicating previous findings and giving novel ones. The HKC provides a valuable resource for exploring eye movement control; the study contrasts the different scenarios of single-sentence and passage reading in hopes of shedding new light on both the universal nature of reading and the unique characteristics of Chinese reading.
This article presents a corpus-based distributional analysis of the usage patterns of a cluster of words and compounds containing the morpheme qia (sic) 'just, exactly', by the aid of an extended concordancer to retrieve representative collocations from their adjacent contexts in Chinese Gigaword. Upon a survey of the historical evolution of the qia cluster with exemplar data and an overview of existing proposals to account for their usages in terms of expectational match, our distributional analysis is conducted to identify the salient collocational or contextual features that lead to a number of interesting findings. Substantial evidences are provided for clarifying the non-word status of qia ru ((sic)(sic)) and qia si ((sic)(sic)) and their similarities, the exchangeability of qiahao ((sic)) and qiaqiao ((sic)), distinct collocational preferences of the adverbs qia (sic), qiaqia ((sic)(sic)) and the others with different subsets of verbs, the prosodic requirement of an even number of syllables for a qia-adverb and its main verb, and the contrastive popularity of qiaqia ((sic)(sic)) vs qiadang ((sic)(sic)) to reveal different usage tendencies between speakers in Taiwan and the Mainland. All these novel findings and insights about the subtle (dis)similarities in the usage and meanings of the qia (sic) cluster suggest that distributional analysis of contextual collocations using large-scale language data remains a powerful tool that can complement other analytical approaches for the advancement of lexical semantic research.
The amount of data has growing significance in exploring cutting-edge materials and a number of datasets have been generated either by hand or automated approaches. However, the materials science field struggles to effectively utilize the abundance of data, especially in applied disciplines where materials are evaluated based on device performance rather than their properties. This article presents a new natural language processing (NLP) task called structured information inference (SII) to address the complexities of information extraction at the device level in materials science. We accomplished this task by tuning GPT-3 on an existing perovskite solar cell FAIR (Findable, Accessible, Interoperable, Reusable) dataset with 91.8% F1-score and extended the dataset with data published since its release. The produced data is formatted and normalized, enabling its direct utilization as input in subsequent data analysis. This feature empowers materials scientists to develop models by selecting high-quality review articles within their domain. Additionally, we designed experiments to predict the electrical performance of solar cells and design materials or devices with targeted parameters using large language models (LLMs). Our results demonstrate comparable performance to traditional machine learning methods without feature selection, highlighting the potential of LLMs to acquire scientific knowledge and design new materials akin to materials scientists.
This chapter introduces the key issues and basic principles of computer(-aided) translation evaluation, covering its historical evolution, context-dependent multi-dimensional nature, and existing methodologies. The major evaluation approaches, including both manual and automatic, are presented with discussion of their strengths and weaknesses.
The scientific literature contains valuable information that can be used for future applications, but manual analysis presents challenges due to its size and disciplinary boundaries. The prevailing solution involves natural language processing (NLP) techniques such as information retrieval. Nonetheless, existing automated systems primarily provide either statistically based shallow information or deep information without traceability, thereby falling short of delivering high-quality and reliable insights. To address this, we propose an innovative approach of leveraging sentiment information embedded within the literature to track the opinions toward materials. In this study, we integrated material knowledge into text representation and constructed opinion data sets to hierarchically train deep learning models, named as Scientific Sentiment Network (SSNet). SSNet can effectively extract knowledge from the energy material literature and accurately categorize expert opinions into challenges and opportunities (94% and 92% accuracy, respectively). By incorporating sentiment features determined by SSNet, we can predict the ranking of emerging thermoelectric materials with a 70% correlation to experimental outcomes. Furthermore, our model achieves a commendable 68% accuracy in predicting suitable nanomaterials for atomic layer deposition (ALD) over time. These promising results offer a practical framework to extract and synthesize knowledge from the scientific literature, thereby accelerating research in the field of nanomaterials.
Haihua Pan (潘海华)合作论文数Department of Linguistics and Modern Languages,The Chinese University of Hong Kong3