This paper extends the stochastic-oracle model of AI-augmented computing to include agentic oracles. Unlike a stationary stochastic oracle, which responds to the same query according to a fixed response distribution across calls, an agentic oracle can pursue a goal autonomously and may access an environment containing task-relevant resources. These capabilities affect both response distributions and token costs beyond what is visible at the query-response interface. We develop a framework for analyzing token costs in Stochastic-Oracle Turing Machines (SOTMs) that compute with agentic oracles. Each call has an orchestration token cost, visible to the caller at the query-response interface, and an agentic token cost, incurred by internal operations not exposed to the caller. We show that an SOTM computing with an agentic oracle that can retain intermediate state can have token-cost advantages over SOTMs using stationary stochastic oracles when solving the same task at the same quality level, both with and without environment access. We also investigate goal-loss risk, including how internal dispatch ordering can reduce exposure to irreversible actions. We provide a goal-loss avoidance criterion, derive progress–retry–goal-loss formulas, establish goal-depth lower bounds on token complexity, characterize token complexity when the probability of goal loss is zero, and show that goal-loss risk can impose an upper bound on the achievable quality of a task involving environment updates.
AI-augmented computing delegates natural language queries, code generation requests, and other open-ended tasks to a cluster of AI models that processes queries and generates responses. This paradigm introduces a resource dimension that neither classical time nor space complexity captures: the cost of sending queries to and receiving responses from such a cluster. We introduce token complexity, a formal resource measure defined as the minimum expected token cost to achieve a specified level of output quality on a task, and develop a taxonomy classifying AI systems by the strength of their probabilistic properties. We develop token complexity within the framework of AI-Oracle Turing machines, in which a probabilistic Turing machine interacts with a stochastic oracle via dedicated query and response tapes. We prove basic theorems establishing that token complexity behaves as expected: monotonicity (higher quality costs more tokens), convexity (quality improvements become progressively more expensive), price sensitivity (small price changes produce bounded cost changes), and price-relativity of task ordering (the token complexity ordering of tasks can reverse depending on the query-to-response cost ratio). We prove that the complexity frontier, defined as the set of all feasible resource bounds in tokens, time, and space, is non-empty, upward-closed, and convex.
The Stochastic-Oracle Turing Machine (SOTM) framework models AI-augmented computation as the interaction of a probabilistic Turing machine with an oracle whose responses are drawn from context-dependent distributions. This paper studies what an SOTM can achieve under two oracle-response schemes: in a cached-response oracle, each distinct query receives one response that is reused on later calls to the same query, while in a fresh-response oracle, each call returns an independent response. In both schemes, the SOTM first computes from its input and internal random source to generate its first query, then proceeds adaptively, computing from its query-response transcript (the record of queries issued and responses received) to generate each subsequent query or produce a final output. Cached responses impose two transcript-based ceilings on achievable performance: a correct-identification ceiling governed by the total variation distance between the transcript distributions induced by the hidden states of the oracle, and an output quality ceiling equal to the expected score of the best output the SOTM can compute from the transcript. Fresh responses can raise these ceilings by allowing repeated calls to accumulate independent evidence toward correct or high-quality outputs. In the binary single-informative-query case, the error probability decreases exponentially in the number of calls to the same query at the Chernoff rate. For output quality, query-count bounds characterize threshold stopping when the score function is incorporated as part of the SOTM, and majority-based amplification bounds characterize the binary candidate-output model when it is not. Together, the results identify how response reuse, transcript information, and access to the score function determine what an SOTM can compute and at what token cost.
Wang introduced the Stochastic-Oracle Turing Machine (SOTM) framework and defined token complexity as the minimum expected cost of interacting with a stochastic oracle needed to attain a specified solution quality for a task. This paper develops an analogous notion for certifying the reliability of a stochastic oracle on a given domain. Certification token complexity is the minimum expected token cost required, with controlled error probability, to distinguish oracles that meet a target reliability level from those that fall below a lower reliability threshold. We construct an SPRT-based certification SOTM that queries the oracle, computes binary correctness scores, and stops when the accumulated log-likelihood evidence crosses a decision threshold. The SOTM halts almost surely, satisfies the desired two-sided error guarantee over the reliability regions to be certified, and yields an explicit upper bound on certification token complexity in terms of the reliability thresholds, the error bound, and the expected per-turn token cost. We then establish a matching information-theoretic lower bound: even with adaptive queries, every error-bounded certification SOTM must incur the same leading-order expected token cost as the SPRT-based construction as the prescribed error bound tends to zero. Together, these bounds characterize the leading-order certification token complexity in the small-error regime.
Parent WeChat groups, established and organized based on technological advancements, serve as crucial platforms for home-school communication. However, various phenomena arising from group pressure within these groups occasionally occur. Through observing participation in parent WeChat groups and conducting in-depth interviews with some parents, the author explores behaviors such as activity, following, and silence exhibited by parents in the "WeChat sign-up chain" phenomenon from the perspective of group pressure. The author analyzes issues arising in parent groups and proposes corresponding strategies.
We introduce a method called AI-WordNet for expanding WordNet using only glosses. Given a lemma and its gloss, AI-WordNet predicts the corresponding lexname for the gloss, decides whether a new synset should be created for the lemma, finds the best hypernym for the new synset, and updates the corresponding hypernym-hyponym relationships. We demonstrate a low-cost implementation of AI-WordNet for noun lemmas using PWN 3.0 as the training dataset. Intrinsic evaluations show that AI-WordNet achieves a high F1 score of 94.6% for lexname prediction, and 64.8% direct hits on true hypernyms with an average distance of 1.374 from the predicted hypernyms to the true hypernyms. We apply AI-WordNet to generate distractors for cloze-question creation on answer keys containing new lemmas not yet included in PWN 3.0, using glosses extracted from Wiktionary. Extrinsic evaluations confirm the high quality of the generated distractors.
This demo introduces Doenba Edit, a user-friendly, AI-powered platform developed by Librum Technologies, Inc., designed for seamless editing and typesetting of academic writing within a word processor-like interface. It supports the entire academic writing workflow, from outlining and idea development to drafting, revising, typesetting, and cross-referencing, assisted by integrated AI tools at every stage. This offers a comprehensive solution for enhancing both the quality and efficiency of producing scholarly work.
Extractive reading comprehension systems are designed to locate the correct answer to a question within a given text. However, a persistent challenge lies in ensuring these models maintain high accuracy in answering questions while reliably recognizing unanswerable queries. Despite significant advances in large language models (LLMs) for reading comprehension, this issue remains critical, particularly as the length of supported contexts continues to expand. To address this challenge, we propose an innovative data augmentation methodology grounded in a multi-agent collaborative framework. Unlike traditional methods, such as the costly human annotation process required for datasets like SQuAD 2.0, our method autonomously generates evidence-based question-answer pairs and systematically constructs unanswerable questions. Using this methodology, we developed the FactGuard-Bench dataset, which comprises 25,220 examples of both answerable and unanswerable question scenarios, with context lengths ranging from 8K to 128K. Experimental evaluations conducted on seven popular LLMs reveal that even the most advanced models achieve only 61.79 importance of a model's ability to reason about unanswerable questions to avoid generating plausible but incorrect answers. By implementing efficient data selection and generation within the multi-agent collaborative framework, our method significantly reduces the traditionally high costs associated with manual annotation and provides valuable insights for the training and optimization of LLMs.
CncRNAs (coding and noncoding RNAs) are a class of bifunctional RNAs that that has both coding and noncoding biological activity. An increasing number of cncRNAs are being identified, prompting reassessment of our knowledge of RNA. However, most existing RNA classification tools are based on binary classification models which are not effective in distinguishing cncRNAs from mRNAs or long noncoding RNAs (lncRNAs). Our statistical analysis demonstrated that mRNA-derived cncRNAs (untranslated mRNAs, untr-mRNAs) and lncRNA-derived cncRNAs (translated ncRNAs, tr-ncRNAs) do not fall in the same cluster. Therefore, in this study, we devised a novel tetra-class RNA classification model that is systematically optimized for RNA feature extraction. According to our model, all human RNAs can be reclassified into one of four categories — mRNA, untr-mRNA, lncRNA, and tr-ncRNA — representing a novel RNA classification system and allowing the discovery of more potential cncRNAs. Further analysis revealed significant differences among the four types of RNAs in tissue-specific expression, functional annotation, sequence composition, and other factors, providing insights into their divergent evolution trajectories. Moreover, investigation of the small tr-ncRNA peptides demonstrated that their evolution is coordinated with that of the the conserved functional small RNAs associated with them. All analysis results have been integrated into a database — TetraRNADB accessible online (http://tetrarnadb.liu-lab.com/).
We introduce the Overall Performance Index (OPI), an intrinsic metric to evaluate retrieval-augmented generation (RAG) mechanisms for applications involving deep-logic queries. OPI is computed as the harmonic mean of two key metrics: the Logical-Relation Correctness Ratio and the average of BERT embedding similarity scores between ground-truth and generated answers. We apply OPI to assess the performance of LangChain, a popular RAG tool, using a logical relations classifier fine-tuned from GPT-4o on the RAG-Dataset-12000 from Hugging Face. Our findings show a strong correlation between BERT embedding similarity scores and extrinsic evaluation scores. Among the commonly used retrievers, the cosine similarity retriever using BERT-based embeddings outperforms others, while the Euclidean distance-based retriever exhibits the weakest performance. Furthermore, we demonstrate that combining multiple retrievers, either algorithmically or by merging retrieved sentences, yields superior performance compared to using any single retriever alone.
Download This Paper Open PDF in Browser Add Paper to My Library Share: Permalink Using these links will ensure access to this page indefinitely Copy URL Copy DOI
Existing tools to detect text generated by a large language model (LLM) have met with certain success, but their performance can drop when dealing with texts in new domains. To tackle this issue, we train a ranking classifier called RoBERTa-Ranker, a modified version of RoBERTa, as a baseline model using a dataset we constructed that includes a wider variety of texts written by humans and generated by various LLMs. We then present a method to fine-tune RoBERTa-Ranker that requires only a small amount of labeled data in a new domain. Experiments show that this fine-tuned domain-aware model outperforms the popular DetectGPT and GPTZero on both in-domain and cross-domain texts, where AI-generated texts may either be in a different domain or generated by a different LLM not used to generate the training datasets. This approach makes it feasible and economical to build a single system to detect AI-generated texts across various domains.
Generative text steganography has received considerable attention in the covert communication community for the benefit of sending secret messages without the need to modify carriers. Existing methods typically choose the next word when generating a stego-text based on conditional probability encoding of candidates, which may lead to generating inadequate words for the underlying secret message. How to generate a semantically controllable stego-text with a high capacity on secure embedding of a secret message is a main challenge. We address this challenge by proposing a new paradigm to generative text steganography that takes advantage of certain social media through apparently normal behaviors from the sender. In particular, we make use of the live commenting feature provided by public video sharing platforms (PVSPs), which allow viewers to make comments on video scenes that will fly on screens when the scenes are shown. We show that this feature can be used to construct a generative steganographic system. The sender generates at random a number of distracting words and a certain invertible matrix called W- d matrix based on the total number of message words and distracting words. The sender then transforms a sequence of indexes of these words to a sequence, selects one or more videos with a sufficiently large number of total frames, and generates a comment on each frame in the sequence. The receiver extracts commented frame indexes, uses the shared stego-key to generate the same W- d matrix as the sender, and obtains the secret message using the inverse of the matrix. The stego-key consists of a vocabulary generator and a W- d matrix generator (WMG) based on pseudorandomly generated numbers. To generate comments on frames that conform to comments made by viewers, we devise a neural ResNet-LSTM model to generate a comment for an input image based on its content. Theoretical analysis shows that commented video frames (CVF) is covert, secure, efficient, and feasible to conceal any message of arbitrary length. We implement CVF and present evaluation results from multiple aspects that our work outperforms the existing stego-methods.
Automatic extraction of legal elements (LEs) from legal documents is essential for constructing a smart court of law. While this task has met with certain success, how to distinguish similar LEs that appear infrequently in the training set remains a challenge, particularly when the number of training documents is limited and the cost of labeling additional documents is high. To overcome this obstacle, we present L-RACL, a neural-net model with label recross attention and contrastive loss to distinguish similar LE labels and capture correlations between LE labels. Extensive experiments show that L-RACL outperforms existing methods. Finally, we finetune ChatGPT to perform LE extraction and show that L-RACL surpasses it.
We propose a CNN-BiLSTM-Attention classifier to classify online short messages in Chinese posted by users on government web portals, so that a message can be directed to one or more government offices. Our model leverages every bit of information to carry out multi-label classification, to make use of different hierarchical text features and the labels information. In particular, our designed method extracts label meaning, the CNN layer extracts local semantic features of the texts, the BiLSTM layer fuses the contextual features of the texts and the local semantic features, and the attention layer selects the most relevant features for each label. We evaluate our model on two public large corpuses, and our high-quality handcraft e-government multi-label dataset, which is constructed by the text annotation tool doccano and consists of 29920 data points. Experimental results show that our proposed method is effective under common multi-label evaluation metrics, achieving micro-f1 of 77.22%, 84.42%, 87.52%, and marco-f1 of 77.68%, 73.37%, 83.57% on these three datasets respectively, confirming that our classifier is robust. We conduct ablation study to evaluate our label embedding method and attention mechanism. Moreover, case study on our handcraft e-government multi-label dataset verifies that our model integrates all types of semantic information of short messages based on different labels to achieve text classification.
We explore how to capture the significance of a sub-text block in an article and how it may be used for text mining tasks. A sub-text block is a sub-sequence of sentences in the article. We formulate the notion of content significance distribution (CSD) of sub-text blocks, referred to as CSD of the first kind and denoted by CSD-1. In particular, we leverage Hugging Face's SentenceTransformer to generate contextual sentence embeddings, and use MoverScore over text embeddings to measure how similar a sub-text block is to the entire text. To overcome the exponential blowup on the number of sub-text blocks, we present an approximation algorithm and show that the approximated CSD-1 is almost identical to the exact CSD-1. Under this approximation, we show that the average and median CSD-1's for news, scholarly research, argument, and narrative articles share the same pattern. We also show that under a certain linear transformation, the complement of the cumulative distribution function of the beta distribution with certain values of $\alpha$ and $\beta$ resembles a CSD-1 curve. We then use CSD-1's to extract linguistic features to train an SVC classifier for assessing how well an article is organized. Through experiments, we show that this method achieves high accuracy for assessing student essays. Moreover, we study CSD of sentence locations, referred to as CSD of the second kind and denoted by CSD-2, and show that average CSD-2's for different types of articles possess distinctive patterns, which either conform common perceptions of article structures or provide rectification with minor deviation.
We present a generative method called CQG for constructing cloze questions from a given article using neural networks and WordNet, with an emphasis on generating multigram distractors. Built on sense disambiguation, text-to-text transformation, WordNet's synset taxonomies and lexical labels, CQG selects an answer key for a given sentence, segments it into a sequence of instances, generates instance-level distractor candidates (IDCs) using a transformer and sibling synsets. It then removes inappropriate IDCs, ranks the remaining IDCs based on contextual embedding similarities, as well as synset and lexical relatedness, forms distractor candidates by combinatorially replacing instances with the corresponding top-ranked IDCs, and checks if they are legitimate phrases. Finally, it selects top-ranked distractor candidates based on contextual semantic similarities to the answer key. Experiments show that this method significantly outperforms SOTA results. Human judges also confirm the high qualities of the generated distractors.
Detecting if two functions in different compiled forms are similar has a wide range of applications in software security. We present a method that leverages both semantic and structural features of functions, learned by a neural-net model on the underlying control-flow graphs (CFGs). In particular, we devise a neural function-similarity regressor (NFSR) with attentions on dual CFGs. We train and evaluate NFSR on a dataset consisting of nearly 4 million functions from over 14 900 binary files. Experiments show that NFSR is superior to the SOTA models of SAFE, Gemini and GMN, especially for binary functions with large CFGs. An ablation study shows that attention on dual CFGs plays a significant role in detecting function similarities.