Large Language Model (LLM)-based agents can solve complex procedural tasks by interacting with environments over multiple turns, but this ability typically depends on large models, long contexts, and repeated inference calls. This makes advanced memory-augmented agents difficult to deploy on resource-constrained devices. We introduce DuoMem, a dual-space distillation framework that transfers procedural problem-solving ability from a large teacher model to compact student models. DuoMem distils in two complementary spaces: (1)context-space distillation, which replaces student-generated memories with higher-quality teacher-generated procedural memories prepended to the student's input, and (2)parameter-space distillation, which fine-tunes lightweight LoRA adapters on successful teacher trajectories. Evaluated on ALFWorld, a challenging embodied decision-making benchmark, DuoMem boosts a 4B-parameter model from 4.3
Large Language Models (LLMs) have demonstrated remarkable capabilities in comprehending and analyzing lengthy sequential inputs, owing to their extensive context windows that allow processing millions of tokens in a single forward pass. However, this paper uncovers a surprising limitation: LLMs fall short when handling long input sequences. We investigate this issue using three datasets and two tasks (sentiment analysis and news categorization) across various LLMs, including Claude 3, Gemini Pro, GPT 3.5 Turbo, Llama 3 Instruct, and Mistral Instruct models. To address this limitation, we propose and evaluate ad-hoc solutions that substantially enhance LLMs' performance on long input sequences by up to 50 and 50
Test-time Reinforcement Learning (TTRL) has shown promise in adapting foundation models for complex tasks at test-time, resulting in large performance improvements. TTRL leverages an elegant two-phase sampling strategy: first, multi-sampling derives a pseudo-label via majority voting, while subsequent downsampling and reward-based fine-tuning encourages the model to explore and learn diverse valid solutions, with the pseudo-label modulating the reward signal. Meanwhile, in-context learning has been widely explored at inference time and demonstrated the ability to enhance model performance without weight updates. However, TTRL's two-phase sampling strategy under-utilizes contextual guidance, which can potentially improve pseudo-label accuracy in the initial exploitation phase while regulating exploration in the second. To address this, we propose context-guided TTRL (CG-TTRL), integrating context dynamically into both sampling phases and propose a method for efficient context selection for on-device applications. Our evaluations on mathematical and scientific QA benchmarks show CG-TTRL outperforms TTRL (e.g. additional 7
Standards developed by the Internet Engineering Task Force (IETF) are critical to the operation of the Internet. Understanding which organisations employ or support participants in the IETF is therefore critical to understanding the development of the Internet. In this paper, we present a longitudinal analysis of affiliation changes in the period 2001-2023, covering 73,764 individuals affiliated with 6,940 organisations. We show that organisational diversity first grew, before remaining broadly constant or even declining (e.g., meeting attendance and RFC authorship). Affiliations have gradually shifted away from North America-based organisations, with an increase in European and Asian organisations. We also observe that a change in affiliation has a short-lived positive impact on participant output and engagement, and increases the chance that an RFC is published, but slows its publication.
Microblogging is a crucial mode of online communication. However, launching a new microblogging platform remains challenging, largely due to network effects. This has resulted in entrenched (and undesirable) dominance by established players, such as X/Twitter. To overcome these network effects, Bluesky, an emerging microblogging platform, introduced starter packs — curated lists of accounts that users can follow with a single click. We ask if starter packs have the potential to tackle the critical problem of social bootstrapping in new online social networks. We assess whether starter packs have indeed been helpful in supporting Bluesky growth. Our dataset includes 25.05 × 10⁶ users and 335.42 × 10³ starter packs with 1.73 × 10⁶ members, covering the entire lifecycle of Bluesky. We study the usage of these starter packs, their ability to drive network and activity growth, and their potential downsides. We also quantify the benefits of starter packs for members and creators on user visibility and activity while identifying potential challenges. By evaluating starter packs’ effectiveness and limitations, we contribute to the broader discourse on platform growth strategies and competitive innovation in the social media landscape.
The Fediverse, a group of interconnected servers providing a variety of interoperable services (e.g. micro-blogging in Mastodon) has gained rapid popularity. This sudden growth, partly driven by Elon Musk's acquisition of Twitter, has created challenges for administrators though. This paper focuses on one particular challenge: content moderation, e.g. the need to remove spam or hate speech. While centralized platforms like Facebook and Twitter rely on automated tools for moderation, their dependence on massive labeled datasets and specialized infrastructure renders them impractical for decentralized, low-resource settings like the Fediverse. In this work, we design and evaluate FedMod, a collaborative content moderation system based on federated learning. Our system enables servers to exchange parameters of partially trained local content moderation models with similar servers, creating a federated model shared among collaborating servers. FedMod demonstrates robust performance on three different content moderation tasks: harmful content detection, bot content detection, and content warning assignment, achieving average per-server macro-F1 scores of 0.71, 0.73, and 0.58, respectively.
Multimodal out-of-context news is a type of misinformation in which the image is used outside of its original context. Many existing works have leveraged multimodal large language models (MLLMs) for detecting out-of-context news. However, observing the limited zero-shot performance of smaller MLLMs, they generally require label-rich fine-tuning and/or expensive API calls to GPT models to improve the performance, which is impractical in low-resource scenarios. In contrast, we aim to improve the performance of small MLLMs in a more label-efficient and cost-effective manner. To this end, we first prompt multiple teacher MLLMs to generate both label predictions and corresponding rationales, which collectively serve as the teachers' knowledge. We then introduce a two-stage knowledge distillation framework to transfer this knowledge to a student MLLM. In Stage 1, we apply LoRA fine-tuning to the student model using all training data. In Stage 2, we further fine-tune the student model using both LoRA fine-tuning and DPO on the data points where teachers' predictions conflict. This two-stage strategy reduces annotation costs and helps the student model uncover subtle patterns in more challenging cases. Experimental results demonstrate that our approach achieves state-of-the-art performance using less than 10
Online comments within news articles are a key way people share opinions. Discovering insightful comments can, however, be challenging for readers. A solution to this problem is using comment curation, whereby professional editors select the highest quality comments manually --- referred to as ''editor-picks''. This paper studies the growing use of professional editor-curation for user-generated comments. We focus on the New York Times as a case study, using a dataset covering 80k articles. We study the characteristics of editor-pick comments, highlighting how editor criteria vary across news sections (e.g. sports, entertainment). We find that editor-pick comments tend to be longer, more relevant to the article, positive in sentiment, and contain low toxicity. Our analysis further reveals that editors within different news sections exhibit differing criteria when they perform comment selection. Thus, we finally propose a set of models that can automatically identify good candidate editor-picks. Our ultimate goal is to reduce editor and journalistic workload, increasing productivity and the quality of curated comments.
How similar are politicians to those who vote for them? This is a critical question at the heart of democratic representation and particularly relevant at times when political dissatisfaction and populism are on the rise. To answer this question we compare the online discourse of elected politicians and their constituents. We collect a two and a half years (September 2020 - February 2023) constituency-level dataset for USA and UK that includes: (i) the Twitter timelines (5.6 Million tweets) of elected political representatives (595 UK Members of Parliament and 433 USA Representatives), (ii) the Nextdoor posts (21.8 Million posts) of the constituency (98.4 constituencies). We find that elected politicians tend to be equally similar to their constituents in terms of content and style regardless of whether a constituency elects a right or left-wing politician. The size of the electoral victory and the level of income of a constituency shows a nuanced picture. The narrower the electoral victory, the more similar the style and the more dissimilar the content is. The lower the income of a constituency, the more similar the content is. In terms of style, poorer constituencies tend to have a more similar sentiment and more dissimilar psychological text traits (i.e. measured with LIWC categories).
Mineral-associated organic carbon (MAOC) constitutes a major fraction of global soil carbon and is assumed less sensitive to climate than particulate organic carbon (POC) due to protection by minerals. Despite its importance for long-term carbon storage, the response of MAOC to changing climates in drylands, which cover more than 40% of the global land area, remains unexplored. Here we assess topsoil organic carbon fractions across global drylands using a standardized field survey in 326 plots from 25 countries and 6 continents. We find that soil biogeochemistry explained the majority of variation in both MAOC and POC. Both carbon fractions decreased with increases in mean annual temperature and reductions in precipitation, with MAOC responding similarly to POC. Therefore, our results suggest that ongoing climate warming and aridification may result in unforeseen carbon losses across global drylands, and that the protective role of minerals may not dampen these effects..
An important concept in organisational behaviour is how hierarchy affects the voice of individuals, whereby members of a given organisation exhibit differing power relations based on their hierarchical position. Although there have been prior studies of the relationship between hierarchy and voice, they tend to focus on more qualitative small-scale methods and do not account for structural aspects of the organisation. This paper develops large-scale computational techniques utilising temporal network analysis to measure the effect that organisational hierarchy has on communication patterns within an organisation, focusing on the structure of pairwise interactions between individuals. We focus on one major organisation as a case study - the Internet Engineering Task Force (IETF) - a major technical standards development organisation for the Internet. A particularly useful feature of the IETF is a transparent hierarchy, where participants take on explicit roles (e.g. Area Directors, Working Group Chairs). Its processes are also open, so we have visibility into the communication of people at different hierarchy levels over a long time period. We utilise a temporal network dataset of 989,911 email interactions among 23,741 participants to study how hierarchy impacts communication patterns. We show that the middle levels of the IETF are growing in terms of their dominance in communications. Higher levels consistently experience a higher proportion of incoming communication than lower levels, with higher levels initiating more communications too. We find that communication tends to flow "up" the hierarchy more than "down". Finally, we find that communication with higher-levels is associated with future communication more than for lower-levels, which we interpret as "facilitation". We conclude by discussing the implications this has on patterns within the wider IETF and for other organisations.
Organizational responsibilities can give people power but also expose them to scrutiny. This tension leads to divergent predictions about the use of potentially sensitive language: power might license it, while exposure might inhibit it. Analysis of peoples' language use in a large corpus of organizational emails using standardized Linguistic Inquiry and Word Count (LIWC) measures shows a systematic difference in the use of words with potentially sensitive (ethnic, religious, or political) connotations. People in positions of relative power are ~3 times less likely to use sensitive words than people more junior to them. The tendency to avoid potentially sensitive language appears to be independent of whether other people are using sensitive language in the same email exchanges, and also independent of whether these words are used in a sensitive context. These results challenge a stereotype about language use and the exercise of power. They suggest that, in at least some circumstances, the exposure and accountability associated with organizational responsibilities are a more significant influence on how people communicate than social power.
This paper introduces Mastodoner, a command-line tool and Python library aimed at simplifying access to public data on Mastodon, a prominent player in the Fediverse --- a decentralized network of interconnected social media platforms. Mastodoner addresses the challenges posed by Mastodon's decentralized nature by providing a unified interface for data collection, instance discovery, and secure data sharing. Through examples and demonstrations, this paper illustrates Mastodoner's capabilities in facilitating researchers' access to and analysis of public Mastodon data, thus advancing research in decentralized social media analytics. The tool and documentation are available at: https://github.com/harisbinzia/mastodoner.
Perennial plants create productive and biodiverse hotspots, known as fertile islands, beneath their canopies. These hotspots largely determine the structure and functioning of drylands worldwide. Despite their ubiquity, the factors controlling fertile islands under conditions of contrasting grazing by livestock, the most prevalent land use in drylands, remain virtually unknown. Here we evaluated the relative importance of grazing pressure and herbivore type, climate and plant functional traits on 24 soil physical and chemical attributes that represent proxies of key ecosystem services related to decomposition, soil fertility, and soil and water conservation. To do this, we conducted a standardized global survey of 288 plots at 88 sites in 25 countries worldwide. We show that aridity and plant traits are the major factors associated with the magnitude of plant effects on fertile islands in grazed drylands worldwide. Grazing pressure had little influence on the capacity of plants to support fertile islands. Taller and wider shrubs and grasses supported stronger island effects. Stable and functional soils tended to be linked to species-rich sites with taller plants. Together, our findings dispel the notion that grazing pressure or herbivore type are linked to the formation or intensification of fertile islands in drylands. Rather, our study suggests that changes in aridity, and processes that alter island identity and therefore plant traits, will have marked effects on how perennial plants support and maintain the functioning of drylands in a more arid and grazed world. In global drylands, soils tend to be more fertile beneath tree, shrub and grass islands. Soil fertility was greater beneath taller and wider plants but was unaffected by either grazing pressure or the type of herbivore.
The advent of regulation, such as the Digital Markets Act, will foster greater interoperability across competing digital platforms. In such regulatory environments, decentralized platforms like Mastodon have pioneered the principles of social data portability. Such platforms are composed of thousands of independent servers, each of which hosts their own social community. To enable transparent interoperability, users can easily migrate their accounts from one server provider to another. In this paper, we examine 8,745 users who switch their server instances in Mastodon. We use this as a case study to examine account portability behavior more broadly. We explore the factors that affect users' decision to switch instances, as well as the impact of switching on their social media engagement and discussion topics. This leads us to build a classifier to show that switching is predictable, with an F1 score of 0.891. We argue that Mastodon serves as an early exemplar of a social media platform that advocates account interoperability and portability. We hope that this study can bring unique insights to a wider and open digital world in the future.
The centralization of web services has raised concerns about critical single points of failure, such as content hosting, name resolution, and certification. To address these issues, the "Decentralized Web" movement advocates for de-centralized alternatives. Distributed Hash Tables (DHTs) have emerged as a key component facilitating this movement, as they offer efficient key/value indexing. The InterPlanetary File System (IPFS) exemplifies this approach by leveraging DHTs for data indexing and distribution. A critical finding of previous studies is that DHT PUT performance for record storage is unacceptably slow, sometimes taking minutes to complete and hindering the adoption of delay-intolerant applications. To address this challenge, this research paper presents three significant contributions. First, we present the design of Optimistic Provide, an approach to accelerate DHT PUT operations in Kademlia-based IPFS networks while maintaining full backward compatibility. Second, we implement and deploy the mechanism and see its usage in the de-facto IPFS deployment, Kubo. Third, we evaluate its effectiveness in the IPFS and Filecoin DHTs. We confirm that we enable sub-second record storage from North America and Europe for 90% of PUT operations while reducing networking overhead by over 40% and maintaining record availability.
The pitfalls of centralized social networks, such as Facebook and Twitter/X, have led to concerns about control, transparency, and accountability. Decentralized social networks have emerged as a result with the goal of empowering users. These decentralized approaches come with their own trade-offs, and therefore multiple architectures exist. In this paper, we conduct the first large-scale analysis of Bluesky, a prominent decentralized microblogging platform. In contrast to alternative approaches (e.g. Mastodon), Bluesky decomposes and opens the key functions of the platform into subcomponents that can be provided by third party stakeholders. We collect a comprehensive dataset covering all the key elements of Bluesky, study user activity and assess the diversity of providers for each sub-components.
The InterPlanetary File System (IPFS) is one of the largest platforms in the growing "Decentralized Web". The increasing popularity of IPFS has attracted large volumes of users and content. Unfortunately, some of this content could be considered "problematic". Content moderation is always hard. With a completely decentralized infrastructure and administration, content moderation in IPFS is even more difficult. In this paper, we examine this challenge. We identify, characterize, and measure the presence of problematic content in IPFS (e.g. subject to takedown notices). Our analysis covers 368,762 files. We analyze the complete content moderation process including how these files are flagged, who hosts and retrieves them. We also measure the efficacy of the process. We analyze content submitted to denylist, showing that notable volumes of problematic content are served, and the lack of a centralized approach facilitates its spread. While we identify fast reactions to takedown requests, we also test the resilience of multiple gateways and show that existing means to filter problematic content can be circumvented. We end by proposing improvements to content moderation that result in 227% increase in the detection of phishing content and reduce the average time to filter such content by 43%.
The recent development of decentralised and interoperable social networks (such as the "fediverse") creates new challenges for content moderators. This is because millions of posts generated on one server can easily "spread" to another, even if the recipient server has very different moderation policies. An obvious solution would be to leverage moderation tools to automatically tag (and filter) posts that contravene moderation policies, e.g. related to toxic speech. Recent work has exploited the conversational context of a post to improve this automatic tagging, e.g. using the replies to a post to help classify if it contains toxic speech. This has shown particular potential in environments with large training sets that contain complete conversations. This, however, creates challenges in a decentralised context, as a single conversation may be fragmented across multiple servers. Thus, each server only has a partial view of an entire conversation because conversations are often federated across servers in a non-synchronized fashion. To address this, we propose a decentralised conversation-aware content moderation approach suitable for the fediverse. Our approach employs a graph deep learning model (GraphNLI) trained locally on each server. The model exploits local data to train a model that combines post and conversational information captured through random walks to detect toxicity. We evaluate our approach with data from Pleroma, a major decentralised and interoperable micro-blogging network containing 2 million conversations. Our model effectively detects toxicity on larger instances, exclusively trained using their local post information (0.8837 macro-F1). Our approach has considerable scope to improve moderation in decentralised and interoperable social networks such as Pleroma or Mastodon.