Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithfully, while fixed canvas decoding must choose target length before denoising. Existing masked diffusion decoding work mainly studies token unmasking order, leaving this length decision under-explored despite its direct effect on coverage and redundancy. We introduce Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entropy from all-mask forward passes and selects the canvas the backbone is most prepared to fill. Relative to a baseline using training corpus length statistics, EV recovers 64.9
Combining keystroke logging, screen recordings, interviews, and text quality assessment in two mixed-methods studies with technical writers, this research (1) identifies defining variables of technical writing processes and (2) examines their correlations with and predictive power for text quality. Study 1, an exploratory investigation with 10 participants, identified 22 distinct writing behaviors under six categories of information searching, information reusing, content shaping, organization structuring, language styling, and layout designing during planning, translating, and reviewing sessions. These behavioral variables, together with time-related variables, were subsequently analyzed as “process indicators” in a comparative experiment with 43 participants across experience levels. Results of Study 2 revealed significant differences among experience levels in writing speed, planning duration, pause, search, reuse, content shaping, and structuring. Detailed planning and systematic content/structure editing were strongly associated with higher-quality texts. Building on these findings, we propose a process model of technical writing, explain its correlations with writing score, and depict process profiles of different experience levels. We also highlight the importance of information processing skills in enhancing writing efficiency, offering empirical guidance for technical writing instruction and professional training.
Technical documentation is written for human developers, but an increasing share of software changes is now authored by autonomous coding agents. Which documents they consult, when, and what follows remain unknown. We conduct a behaviour-grounded study of agent-documentation interaction across two public datasets: 557 agentic coding sessions from SWE-chat, yielding 94,813 development events including 3,033 documentation interactions; and 33,097 agentic pull requests from AIDev, with 690,260 classified file-level change records. Four findings challenge current documentation practice. First, agents' documentation work is dominated by agent-facing artefacts: instruction files and working notes account for 60.5
This study addresses the need for a standardized scale to evaluate the user experience (UX) of technical documentation, highlighting the role of high -quality documentation in enhancing user support in the fast-evolving technological landscape. Employing a multi-faceted approach, the authors reviewed existing literature, conducted interviews with users, and perfOrmed exploratory experiments. This multi-fitceted methodology fiwilitated the identification of key UX dimensions and textual elements within technical documentation, leading to the development of an initial evaluation scale. The scale underwent a series of reliability and validity tests, including cross-validation with existing UX scales and user testing involving documents of varied quality levels. The findings confirmed the effectiveness of the scale in providing nuanced insights into UX of technical documentation. The study successfully generated a reliable and efficient scale tailored for the UX evaluation of technical documentation. This scale enables documentation teams to assess and improve documentation quality based on user fredback, enhancing the overall product experience and potentially boosting market share. By bridging the existing gap in UX evaluation tools specifically designed for technical documentation, the validated scale contributes to the body of knowledge in technical communication.
This research delves into understanding the behaviors and characteristics of Chinese developers in relation to their use of technical documentation, which is crucial for creating high-quality developer documentation. We conducted interviews with 25 software developers and surveyed 177 participants, using the preliminary interview findings to inform the survey design. Our approach encompassed traditional user research methods, including persona and user journey mapping, to develop typical personas and information journeys based on the qualitative data from the interviews and quantitative results from the survey. Our results revealed distinct characteristics and differences between junior and senior developers in terms of their use of technical documentation, broadly categorized into personality traits, learning habits, and working habits. We observed that the information journey of both groups typically encompasses four stages: Exploration, Understanding, Practice, and Application. Consequently, we created two distinct personas and information journey maps to represent these two developer groups. Our findings highlight that developers prioritize the content, organization, and maintenance aspects of documentation. In conclusion, we recommend organizing documentation content to align with developers' information journeys, tailoring documentation to meet the needs of developers at various levels, and focusing on the content, organization, and maintenance aspects of documentation.
This pilot UX study aims to establish a user experience (UX) approach for assessing the quality of developer documentation, specifically for the OceanBase database company. A usability test was designed to examine task completion rates, times, errors, and satisfaction. Additionally, biometric equipment, such as eye-tracking, EEG, facial expression recognition, and measures of visual and mental fatigue, were utilized to analyze participants' experiences. The PANAS scale was employed to collect self-reported emotions. The hybrid evaluation method revealed that the OceanBase database documentation suffers from poor usability, understanding, and findability. The root causes of these problems include weak functionality, subpar interaction design, lack of conciseness and explanation, and inadequate structure and legibility. Participants experienced negative emotions and increased mental fatigue after using the documentation, indicating a substantial cognitive load. These findings will inform future improvements to the documentation and provide a foundational UX model for developer documentation. The UX study process may also be a reference for practitioners or researchers conducting similar research on technical documents.
With the rapid development of artificial intelligence technology and increasing material data, machine learning- and artificial intelligence-assisted design of high-performance steel materials is becoming a mainstream paradigm in materials science. Machine learning methods, based on an interdisciplinary discipline between computer science, statistics and material science, are good at discovering correlations between numerous data points. Compared with the traditional physical modeling method in material science, the main advantage of machine learning is that it overcomes the complex physical mechanisms of the material itself and provides a new perspective for the research and development of novel materials. This review starts with data preprocessing and the introduction of different machine learning models, including algorithm selection and model evaluation. Then, some successful cases of applying machine learning methods in the field of steel research are reviewed based on the main theme of optimizing composition, structure, processing, and performance. The application of machine learning methods to the performance-oriented inverse design of material composition and detection of steel defects is also reviewed. Finally, the applicability and limitations of machine learning in the material field are summarized, and future directions and prospects are discussed.
Purpose: This study updates our understanding of the group features of China's technical communication coming out of the COVID-19 pandemic. Our research uncovers workplace inequities in the profession by identifying and analyzing a wide range of professional differences in knowledge, skills, experience, practice, performance, benefits, opportunities, challenges, and discoveries. It is more than just a diversity report. We seek to help academics and practitioners across the world develop a basic grasp of China's technical communication, practitioners, and working conditions from a diversity, equity, and inclusion (DEI) perspective. Method: We designed a four-part survey with 50 questions to examine DEI variables in several areas such as demographics, professional activities, career development, and challenges and problems. A total of 259 technical communicators from a target population of about 1,200 responded to our questionnaire. Results: Diversity is an intrinsic feature of China's technical communication because of its short history of professionalization. Practitioners' educational backgrounds, language ability, job titles, affiliated departments, working activities and deliverables, and so on all exhibit diversity. Because of the lack of DEI initiatives, many participants reported structural inequalities in their career development. Conclusion: The DEI situation in the field of China's technical communication is incarnated as a collective professional identity crisis in practitioners. This identity crisis has historical, societal, organizational, individual, and environmental reasons. To tackle it, we propose inclusive development as an effective DEI initiative.
为解决高速公路合流区在交通需求较高条件下的交通流不稳定和拥堵问题,以提高合流区通行能力和车辆合流过程协调性以及消除合流冲突为切入点,提出了在全智能网联车辆(CAV)环境下基于编队的协同合流(PBCM)策略.在合流区上游匝道路段设置一定长度的编队区,当编队区内的CAV数量满足一定规模后,即触发PBCM策略的执行.策略共包含匝道编队区CAV和主路CAV初始合流时间计算、匝道和主路CAV编队方案制定以及车队头车轨迹规划算法3部分.利用MATLAB构建了高速公路合流仿真环境,实验结果显示,当总交通需求为3 100辆/h时,PBCM策略通过合流点的车辆数和主路平均速度最高分别比现有基于单个匝道和主路车辆协同(SVBCM)的策略提高了50.7%和20.0%,延误降低46.7%.仿真结果验证了PBCM策略在解决高交通需求条件下合流区拥堵问题的有效性,在匝道合流率较大的情况下仍可以维持主路上游较高的行驶速度.
针对目前智能网联车(connected and autonomous vehicle,CAV)通过交叉口的轨迹规划算法无法兼顾效率与安全协同最优的问题,根据CAV驶入交叉口通信范围时的不同初始行驶状态引入动态距离窗(dynamic dis?tance windows,DDW)概念,提出适用于可控安全行驶条件下通行效率最优的轨迹规划算法.算法根据CAV初始行驶状态参数和信号灯信息、最大舒适加/减速度和道路限速约束条件,获得CAV初始行驶状态对应的DDW.针对CAV初始位置与停车线上游特定位置间的距离处于DDW范围之内和之外的两种情况,分别设计相应的轨迹规划算法,实现CAV通过交叉口延误最小.仿真结果显示,所提出的算法可有效提高CAV通过交叉口的效率,并且具有更小的速度波动及更平滑的时空轨迹.
This paper was presented at the Invited Panel session “Technical Communication in China”. Technical documentation quality is an essential part of product help center quality. As user information behavior gains increasing attention in user-interface design, for help centers, understanding user’s information behavior of documentation provides insights to better develop product resources. To improve help center quality, we focus on typical technical documentation users—developers and their information behavior. We used combined sociological approaches and conducted a two-phase user research: an in-depth interview was first performed to gain basic idea of developers’ documentation acquiring and using habits; then, a questionnaire survey was conducted to validate former result and further explore the factors influence those behaviors. This paper present how developers need, acquire and use technical documentations, and discussed the related factors. According to the result, we found that developers have preference in documentation acquiring methods, besides, quickly locating target information is their common way to start using a documentation. Furthermore, gender and user documentation dependency both have impact on their information behaviors and documentation using experience. In conclusion, developers’ information behaviors of technical documentation have patterns, and these patterns should be carefully considered in help centers’ future development.
This paper was presented at the Invited Panel session “Technical Communication in China”. There has been various research on the reading time and legibility of online texts with people’s tendency to online materials. Text-related attributes like font size or letterspacing are commonly used variables in this field. The objective of this study is to investigate the influential factors on the reading time of Chinese technical documentation, and to build a Decision Tree model to predict its reading time. In the experiment, log data including information of over a million user visits from a cloud service provider’s website are collected. User’s visit time, stay time, visit step, visit device and many other data fields are recorded in a user session. In addition to user behavioral data from log files, data metrics concerning technical documentation itself are also collected. For all documents used in the experiment, their word counts, image counts, link counts and section counts are scraped using web crawlers. The linear correlation analysis is applied in order to explore the correlations between variables for predictions. The results show that a 75 percent accuracy is achieved using the Decision Tree model.
This paper was presented at the Invited Panel session “Technical Communication in China”. Findability is one of the most important qualitative factors of websites. With the rapid growth in navigation complexity and in number of technical documentations in help centers, whether users can easily locate the target document could directly determine the information retrieval task outcome. Providing users with a fine guide to target documents and then helping them find solutions to their problems is the most important function of a help center. Investigation on user search behavior data and perceived findability of documentation has to be done in order to further apply website log data to predicting user subjective assessment. In this paper we analyze the correlation between subjective document findability, subjective task complexity, and user search behavior. We found several search behavior metrics which significantly correlate with the two subjective measures above.
高速公路交通流具有多维时空动态特征,其实际调查的数据存在过量随机异常的数据,导致有关调查数据的质量无法满足高速公路主动安全管理对路网各层级交通流调控的需求.利用张量理论所具有的良好多维时空数据处理能力,构建考虑短期波动、长期趋势和车道空间信息的多维时空动态张量矩阵,提出一种基于多维张量分解思维的梯度下降Tucker分解的数据质量控制算法,有效弥补传统数据质量控制方法对交通流数据内在时空关联信息利用不足的情况.选取G4京港澳高速公路杜家坎路段实际速度数据作为研究对象,选择车道维度、时间维度、时间间隔维度,构建多种不同张量矩阵形式,对算法进行实证分析.结果表明:所提出的数据质量控制算法具有良好的高速公路交通流数据修复效果.其中,以车道、天数、5 min采集时间间隔3个维度所构建的张量形式修复效果最好,95%测试数据的修复值与实际值的误差在(-5%,+5%)范围内.
Help centers are mainly designed to assist users with their product uses. The question as to how we measure the quality of a help center remains unanswered. As the first step of a joint research initiated by Peking University and Baidu Cloud that aims to develop a set of computable metrics to evaluate the quality of help centers, this experience report shares the results of data analysis on correlation between user behavioral data and technical documentation quality. The documents and data we use are a suite of cloud computing services provided by Baidu Cloud. The report begins with an introduction of the research goal; following reviews on the related work, it then lays out the design of the experiments with user data collected from Baidu Cloud. In our experiments, we categorize all documents into three groups and try to identify which metrics would affect documentation quality most. The result shows that the key index that contributes most to the model is PV/UV. At last, the report concludes with our current experimental efforts and future work in our plan.
This paper presents users' preferred line length and its effect to user performance when they read Chinese technical text. A web-shaped software is developed to record user behaviors. The participants are asked to adjust the length of the text until they feel comfortable with. The study includes two tasks, pre-reading task and while-reading task. In pre-reading task, participants adjust the line length and stop at their preferred line length; while-reading task record user's further adjusted line length when they are reading. The study shows that 76 CPL (character per line) is the preferred line length with a good score and task-solving efficiency. This finding can be used to help decide the line length for Chinese technical text.
针对目前短时交通流预测算法多考虑交通流的低维信息特征,导致无法满足预测精准度要求等问题,引入高精度低秩张量填充理论(HALRTC),构建基于周、天、时段等多时间维度的动态张量模型,设计了一种融合高维交通流特征的短时交通流预测算法,并以京港澳高速公路杜家坎路段交通流速度数据为例进行实证验证.研究结果显示,算法能够基于较少历史数据较快达到良好预测效果,可有效实现针对工作日与非工作日的交通流预测,平均绝对误差(MAE)平均值约为3.6%,并能及时跟踪交通流波动性.在缺失数据情况下,所提出算法预测精度随数据缺失比例增大而降低,但相较于3种经典预测算法可表现出更好的预测精度.
For the sake of studying the characteristics of traffic conflicts involving non-motor vehicles at city road intersections,univariate and multivariate regression analyses were conducted between the numbers of straight-go non-motor vehicles,fight-turn motor vehicles,equivalent cars after converting and the number of conflicts of motors with non-motor vehicles per unit time in the observation time according to actual operation characteristics of non-motor vehicles at a typical intersection in Hohhot.In order to reflect the severity of conflicts involving non-motor vehicles comprehensively and accurately,severity models considering and normalizing the influences of distance,angle and speed of conflict on the severity were built.Results show that the quadratic polynomial fitted with the number of straight-go non-motor vehicles and the number of conflicts has the best predictive performance among the 9 univariate regression models and its forecast accuracy of conflicts number is 85.3%,that the function fitted with the number of straight-go nonmotor vehicles,the number of equivalent cars and the number of conflicts has the best predictive performance among the 4 multivariate regression models and its forecast accuracy is 92.4%,so it is the best forecast function of conflicts number,and that the established severity models for non-motor vehicle conflicts can be used to measure the severity of traffic conflict in practice.
Although technology plays a major role in Chinese business and everyday life, the field of technical communication is still largely unknown within Chinese industry and academia. Since the 1990s, attempts have been made to establish technical communication courses and programs in Chinese universities. In this article, we will describe the recent developments at Peking University.