毕马威会计师事务所(英文简称:KPMG,台湾称安侯建业)是世界上最大的专业服务机构之一。该公司雇佣员工约为136,500人, 并且在全球超过140个国家或地区设有分支机构。2008年,毕马威实现227亿美元的收入,比2007年增长14.5%。毕马威提供三类主营服务,分别是:审计、税务和咨询。毕马威是四大国际会计师事务所之一,而其他三大分别是普华永道、德勤和安永。
As LLM agents become capable of increasingly long-horizon tasks, evaluating their performance in economic systems is becoming increasingly important. Unlike existing benchmarks that primarily evaluate a single agent interacting with a passive environment, economic systems are inherently multi-agent, requiring autonomous agents to communicate, negotiate, and transact while pursuing their own objectives over extended periods. We introduce CoffeeBench, a benchmark for evaluating LLM agents in a long-horizon multi-agent economy composed of heterogeneous firms. In CoffeeBench, two farmers, two roasters, and two retailers autonomously operate their businesses over a 90-day simulation, each seeking to maximize cumulative net income through communication and transactions while managing cash, inventory, and pricing. The evaluated model controls one coffee roaster, while the remaining firms are controlled by fixed reference agents. Across several recent open-weight and proprietary LLMs, all models outperform a passive baseline that takes no actions, with most achieving positive net income. Analysis of agent behavior reveals substantial differences in long-horizon economic interaction: higher-performing models communicate more actively with other firms, whereas Claude~Haiku~4.5 exhibits an idle-drift failure mode, repeatedly choosing inaction despite producing coherent assessments and plans. We release our code and agent trajectories to support future research.
We demonstrate the application of a quantum feature extraction method to enhance multi-class image classification for space applications. By harnessing the dynamics of many-body spin Hamiltonians, the method generates expressive quantum features that, when combined with classical processing, lead to quantum-enhanced classification accuracy. Using a strong and well-established ResNet50 baseline, we achieved a maximum classical accuracy of 83
Large language models (LLMs) often generate fluent but unfounded claims, or hallucinations, which fall into two types: (i) faithfulness violations - misusing user context - and (ii) factuality violations - errors from internal knowledge. Proper mitigation depends on knowing whether a model's answer is based on the prompt or its internal weights. This work focuses on the problem of contributive attribution: identifying the dominant knowledge source behind each output. We show that a probe, a simple linear classifier trained on model hidden representations, can reliably predict contributive attribution. For its training, we introduce AttriWiki, a self-supervised data pipeline that prompts models to recall withheld entities from memory or read them from context, generating labelled examples automatically. Probes trained on AttriWiki data reveal a strong attribution signal, achieving up to 0.96 Macro-F1 on Llama-3.1-8B, Mistral-7B, and Qwen-7B, transferring to out-of-domain benchmarks (SQuAD, WebQuestions) with 0.94-0.99 Macro-F1 without retraining. Attribution mismatches raise error rates by up to 70
Low Earth orbit (LEO) satellite systems are becoming an increasingly important component of global communications infrastructure, providing broadband access, enterprise connectivity, and direct-to-device services in competition with terrestrial networks. At the same time, orbital space is a congestible shared resource: satellite deployments increase conjunction risk and debris, imposing external costs on other operators.This paper analyses how these features interact by modelling LEO satellite broadband as a capacity-constrained oligopoly operating under an orbital congestion externality. We develop a two-stage model in which satellite operators first choose constellation size and research and development (R&D) investment, and subsequently compete in quantities subject to binding capacity constraints. Orbital congestion damages depend on aggregate satellite deployment, while operators are privately exposed to only a fraction of the resulting congestion risk.Three results emerge. First, oligopolistic competition and incomplete congestion internalisation generate distinct distortions: output and innovation are inefficiently low due to market power, while satellite deployment is excessive when firms do not face the full marginal social cost of congestion. Second, R&D interacts non-trivially with congestion risk through its effect on throughput per satellite, creating substitution between “more satellites” and “smarter satellites.” Third, regulatory instruments such as Pigouvian satellite charges or tradable conjunction-risk permits can correct deployment incentives but do not eliminate distortions arising from imperfect competition.These results highlight that congestion pricing and competition policy operate as complementary instruments in the governance of emerging satellite broadband markets, with implications for spectrum policy, launch regulation, and the management of shared orbital resources.