宁波宏利集团有限公司,是专业生产出口针织服装和花色针织布的大型企业。至今有20多年的历史。是省内重要的针织品出口基地。拥有自营进出口经营权,2002年7月一次性通过质量体系认证。
Complexly structured data present in documents and web content pose significant challenges for accurate MLLM reasoning. Although MLLMs have advanced substantially, they continue to struggle with intricate data formats such as nested tables and multi-dimensional charts, often leading to hallucinations. This paper explores the capabilities of LLMs and MLLMs in understanding and answering questions from complex data found in PDF documents by leveraging a pre-processing pipeline consisting of industrial and open-source tools. Our results showcase that incorporating RAG and pre-processing tools enables MLLMs to achieve approximately 5 https://github.com/manulife-ai/financialqa .
Contextual bandits are incredibly useful in many practical problems. We go one step further by devising a more realistic problem that combines: (1) contextual bandits with dense arm features, (2) non-linear reward functions, and (3) a generalization of correlated bandits where reward distributions change over time but the degree of correlation maintains. This formulation lends itself to a wider set of applications such as recommendation tasks. To solve this problem, we introduce *conditionally coupled contextual* ($C_3$) Thompson sampling for Bernoulli bandits. It combines an improved Nadaraya-Watson estimator on an embedding space with Thompson sampling that allows online learning without retraining. Empirical results show that $C_3$ outperforms the next best algorithm by 5.7% lower average cumulative regret on four OpenML tabular datasets as well as demonstrating a 12.4% click lift on Microsoft News Dataset (MIND) compared to other algorithms.
Artificial intelligence-based autonomous decision systems are being implemented in the fields of finance, healthcare, transportation, and governance of the people. They are used to improve the accuracy, speed, and scalability of decisions through processing high volumes of data and autonomously generating responses without requiring human intervention. Nevertheless, the adoption of autonomous decision systems is rapid, and it presents new governance issues in transparency, accountability, algorithms, as well as operational risks. The old models of governance still depend on manual supervision and fixed policy enforcement that cannot be used to deal with the dynamic AI-driven environments. The following paper proposes a risk-based governing framework that is aimed at providing responsible and trustworthy functioning of autonomous decision systems. The suggested framework incorporates risk identification, risk assessment, governance policy enforcement, and on-going monitoring mechanisms to handle operational and ethical risks. The framework enables organizations to implement the right governance controls by categorizing autonomous systems according to the level of risk and still providing flexibility in operations. Moreover, the architecture has elements of explainability and compliance monitoring to reinforce transparency and accountability. The research emphasizes the role of risk-based governance in allowing organizations to reduce the risks associated with decisions, enhance the adherence to regulations, and increase trust in AI-based decision systems. The suggested framework offers an organized method of governance that facilitates innovation and sustainable use of autonomous decision technologies.
Prompt optimization can improve multi-agent LLM systems, but the prompts being optimized often serve two entangled roles: generating task-relevant content and specifying execution-critical protocols, such as message routing, output formatting, and termination signals, on which the underlying code relies. As a result, a prompt edit intended to improve content generation can inadvertently corrupt the protocol and cause the entire agent pipeline to fail. Our key observation is that these two roles have different representations: execution protocols are typically structured, while task-relevant content is usually expressed in unstructured language. Based on this, we propose control-data flow separation, where execution-critical control is represented as typed, validated program objects, while task-relevant language remains the optimizable data flow for agent communication. This design allows optimizers to improve multi-agent behavior without exposing the routing or formatting interface to prompt drift. Across synthetic reasoning, collaborative review generation, and insurance rating workflows, our framework empirically achieves 100
Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop reasoning. However, these benchmarks are typically short-horizon, requiring only a small number of retrieval or inference steps, and provide limited evidence of reliability on real-world tasks that involve following manuals spanning hundreds of pages with complex, interdependent guidelines. In this paper, we introduce Tasks over Application Manuals (TAM), a benchmark for evaluating long-horizon procedural reasoning. We construct TAM by curating real-world tasks from two domains: ICD-10-CM clinical coding (mapping medical conditions to diagnostic codes) and U.S. federal sentencing (computing crime sentencing guideline outcomes, specifically offense levels), with human-validated labels. Each task requires following an authoritative manual with tens of thousands of rules and executing a sequence of interdependent steps across different sections to produce an exact answer. We evaluate general-purpose prompting approaches, including retrieval-augmented generation, ReAct-style prompting, and an agent-harness baseline on GPT-5, and find that the best exact-match performance remains extremely low: 1