优步体育用品有限公司是一家专业化生产健身器材的高新技术企业;公司集研发设计、生产和营销推广为一体;产品全部按国际高水平的欧洲EN957安全标准设计、制造;以电动跑步机为主导,涉及健身车,力量训练器,休闲训练器等系列。 优步产品以国际流行的款式,优良的品质以及合理的价格,配合公司良好的服务来满足广大客户及顾客的需求。
Multi-agent LLM pipelines produce contradictory evidence on whether team diversity improves output quality: heterogeneous Mixture-of-Agents teams outperform single models, yet homogeneous Self-MoA teams consistently win under synthesis-based aggregation. We propose a resolution by identifying the selection bottleneck-a crossover threshold in aggregation quality that determines whether diversity helps or hurts. Under this model, we obtain a closed-form crossover threshold s & lowast; (Proposition 1) that separates the regimes where diversity helps and hurts. In a targeted experiment spanning 42 tasks across seven categories (N=210), a diverse team with judge-based selection achieves a win rate of 0.810 against a single-model baseline, while a homogeneous team scores 0.512-near chance (Glass's Delta=2.07). Judge-based selection outperforms MoA-style synthesis by Delta(WR)=+0.631-the synthesis approach is preferred over the baseline in zero of 42 tasks by the judge panel. A decoupled evaluation with fully independent judges confirms all directional findings (Spearman rho=0.90); win rates attenuate by 53-67% under independent evaluation relative to the primary estimates, consistent with partial measurement circularity in the judge-based cells. Exploratory evidence suggests that including a weaker model improves performance while reducing cost (p<10(-4), not pre-registered). Our results suggest that selector quality may be a more impactful design lever than generator diversity in single-round generate-then-select pipelines, with a specific Opus-only singleton as the baseline.
Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D world around them. Existing approaches either inject features from a frozen 3D foundation model without an objective that ensures the policy uses them, or constrain geometry with sparse box and map losses that provide no dense spatial signal. We introduce VLGA, the first vision-language-action model supervised to reconstruct the dense 3D world it drives through. VLGA introduces geometry as a fourth modality alongside vision, language, and action through a dedicated expert supervised by a per-pixel pointmap regression loss against LiDAR. Extensive experiments conducted on challenging nuScenes and Bench2Drive datasets for open-loop and closed-loop evaluations, respectively, show the superiority of VLGA over counterpart VLA methods. In particular, on open-loop nuScenes, VLGA sets a new state of the art among VLA methods without ego status, with the lowest L2 (0.50 m average) and 3-second collision rate (0.18%). On closed-loop Bench2Drive, VLGA attains the state-of-the-art driving score of 79.08, +0.71 over the strongest prior VLA, at comparable efficiency and comfort.
We exploit unexpected corporate data breaches to study the loss and repair of corporate reputation. Reputation loss decreases equity and brand values, increases customer churn, and prompts more negative media coverage. Firms repair their reputation by increasing their charitable donations and have CSR scores that are more than 0.5 standard deviations higher. They increase political contributions, employee wages, and IT investment. These actions are targeted to stakeholders that are particularly important or in situations that are particularly salient to their stakeholders. We observe similar dynamics of reputation loss and repair following the release of negative news about firms' social behaviors.
Mobile applications in large-scale distributed systems are susceptible to backend service failures, yet traditional chaos engineering approaches cannot scale mobile testing due to the combinatorial explosion of flows, locations, and failure scenarios that need validation. We present an automated mobile chaos testing system that integrates DragonCrawl, an LLM-based mobile testing platform, with uHavoc, a service-level fault injection system. The key insight is that adaptive AI-driven test execution can navigate mobile applications under degraded backend conditions, eliminating the need to manually write test cases for each combination of user flow, city, and failure type. Since Q1 2024, our system has executed over 180,000 automated chaos tests across 47 critical flows in Uber's Rider, Driver, and Eats applications, representing approximately 39,000 hours of manual testing effort that would be impractical at this scale. We identified 23 resilience risks, with 70
We present the Agentic AI Detection and Response (ADR) system, the first large-scale, production-proven enterprise framework for securing AI agents operating through the Model Context Protocol (MCP). We identify three persistent challenges in this domain: (1) limited observability – existing Endpoint Detection and Response (EDR) tools see file writes but not the agent reasoning, prompts, or causal chains linking intent to execution; (2) insufficient robustness – static defenses constrained by pre-defined rules fail to generalize across diverse attack techniques and enterprise contexts; and (3) high detection costs – LLM-based inference is prohibitively expensive at scale. ADR addresses these challenges via three components: the ADR Sensor for high-fidelity agentic telemetry, the ADR Explorer for systematic pre-deployment red teaming and hard-example generation, and the ADR Detector for scalable, two-tier online detection combining fast triage with context-aware reasoning. Deployed at Uber for over ten months, ADR has sustained reliable detection in production with growing adoption reaching over 7,200 unique hosts and processing over 10,000 agent sessions daily, uncovering hundreds of credential exposures across 26 categories and enabling a shift-left prevention layer (97.2