简要回顾人工智能发展史及前沿热点.概述人工智能技术在工业各场景下的应用现状,指出工业场景充分发挥人工智能威力的前提条件,提出人工智能首要是优化资源配置效率,探讨人工智能落地工业典型场景的路径.
Offline reinforcement learning (RL) defines the task of learning from a static logged dataset without continually interacting with the environment. The distribution shift between the learned policy and the behavior policy makes it necessary for the value function to stay conservative such that out-of-distribution (OOD) actions will not be severely overestimated. However, existing approaches, penalizing the unseen actions or regularizing with the behavior policy, are too pessimistic, which suppresses the generalization of the value function and hinders the performance improvement. This paper explores mild but enough conservatism for offline learning while not harming generalization. We propose Mildly Conservative Q-learning (MCQ), where OOD actions are actively trained by assigning them proper pseudo Q values. We theoretically show that MCQ induces a policy that behaves at least as well as the behavior policy and no erroneous overestimation will occur for OOD actions. Experimental results on the D4RL benchmarks demonstrate that MCQ achieves remarkable performance compared with prior work. Furthermore, MCQ shows superior generalization ability when transferring from offline to online, and significantly outperforms baselines. Our code is publicly available at https://github.com/dmksjfl/MCQ.
在疾病诊断、手术引导及放射性治疗等图像辅助诊疗场景中,将不同时间、不同模态或不同设备的图像通过合理的空间变换进行配准是必要的处理流程之一.随着深度学习的快速发展,基于深度学习的医学图像配准研究以其耗时短、精度高的优势吸引了研究者的广泛关注.本文全面整理了2015—2019年深度医学图像配准方向的论文,系统地分析了深度医学图像配准领域的最新研究进展,展现了深度配准算法研究从迭代优化到一步预测、从有监督学习到无监督学习的总体发展趋势.具体来说,本文在界定深度医学图像配准问题和介绍配准研究分类方法的基础上,以相关算法的网络训练过程中所使用的监督信息多少作为分类标准,将深度医学图像配准划分为全监督、双监督与弱监督、无监督医学图像配准方法.全监督配准方法通过采用随机变换、传统算法和模型生成等方式获取近似的金标准作为监督信息;双监督、无监督配准方法通过引入图像相似度损失、标签相似度损失等其他监督信息以降低对金标准的依赖;无监督配准方法则完全消除对标注数据的需要,仅使用图像相似度损失和正则化损失监督网络训练.目前,无监督医学图像算法已经成为医学图像配准领域的研究重点,在无需获得代价高昂的标注信息下就能够取得与有监督和传统方法相当甚至更高的配准精度.在此基础上,本文进一步讨论了医学图像配准研究后续可能的4个未来挑战,希望能够为更高精度、更高效率的深度医学图像配准算法的研究提供方向,并推动深度医学图像配准技术在临床诊疗中落地应用.
因受到光线散射和吸收、水体杂质、人工光源等因素影响,水下成像质量较低,很难满足生产作业的需求,而水下图像的增强和复原技术有助于提升水下机器视觉的能力.为帮助研究者掌握水下图像处理领域的研究方法和现有技术,对水下图像增强和复原方法进行综述.首先对水下图像存在的主要退化类型进行分析;分别对水下图像增强、复原的经典方法和最新进展进行总结,系统梳理了水下图像质量评测体系和公开数据集;最后对水下图像处理未来的研究趋势进行了展望.
近年来,强化学习在游戏、机器人控制等序列决策领域都获得了巨大的成功,但是大量实际问题中奖励信号十分稀疏,导致智能体难以从与环境的交互中学习到最优的策略,这一问题被称为稀疏奖励问题.稀疏奖励问题的研究能够促进强化学习实际应用与落地,在强化学习理论研究中具有重要意义.本文调研了稀疏奖励问题的研究现状,以外部引导信息为线索,分别介绍了奖励塑造、模仿学习、课程学习、事后经验回放、好奇心驱动、分层强化学习等方法.本文在稀疏奖励环境Fetch Reach上实现了以上6类方法的代表性算法进行实验验证和比较分析.使用外部引导信息的算法平均表现好于无外部引导信息的算法,但是后者对数据的依赖性更低,两类方法均具有重要的研究意义.最后,本文对稀疏奖励算法研究进行了总结与展望.
Lack of existing research realize how to share critical resource cooperation in crowdsourcing model ,this paper at-tempts to analysis the value capture theory and the classical model of business models ,extracting important influencing factors ,build the crow dsourcing model based on value identity key resources sharing theory model together .On this ba-sis ,in the view of the party aw arding contract and the bag special party development scale ,real participation crow dsourc-ing project oriented network user experience questionnaire ,using of structural equation model of the party and meet smart-PLS3 .0 package party structure equation model to simulate ,the bootstrapping iteration .Studies have show n that the level of cooperation crowdsourcing and collaborative sharing of resources was a significant positive correlation ,while the degree of cooperation and the value of crowdsourcing identity was a significant positive correlation between the value of collabora-tive sharing of identity and resources was a significant positive correlation between the value of identity for the public pack the level of cooperation and collaborative sharing of resources incomplete mediating variables ,market competition intensity has no direct impact on the regulation of the level of cooperation crow dsourcing and collaborative sharing of resources .
Because of the worldwide rapid development of MOOC, academic researches and industrial applications of MOOC have become a branch of the major concerns in modern education and information technology fields. This paper focuses on low completion phenomenon in MOOC environment and proposes an explicable approach to find out hidden reasons convincingly. Different from existing works, this approach utilizes data mining methods to make quantitative analysis. It employs learners clustering basing on their study features at first, aiming at discovering inactive learners automatically. These learners are representative of low completion in course study on MOOC platform. Their study behaviors and interactions with website are analyzed with association rules mining in order to explore potential patterns and rules. The extracted rules are used to find out and explain the reasons for low completion in MOOC environment. The experimental result on practical XuetangX platform reveals several strong rules with high support, confidence and lift, which can be regarded as evidence and reference for further explanations of reasons for low completion in MOOC environment.
Domain ontology is a collection of domain-specific concepts and their interrelationships, which provide an abstract view of the application domain and is used in many areas such as semantic mining(SM) and natural language processing(NLP). But the direct construction of Domain ontology manually is labor intensive and time consuming, while auto-generated Domain-specific Lexical Repository can be used to build domain ontology as an indispensable component. In this paper, we propose a two-stage method to build domain-specific lexical repository making use of the dump service of Chinese Wikipedia. The main idea is that only concepts strongly semantic-related to the multi roots we choose are incorporate into the repository. First we use the dump service for all pages(zhwiki-all-pages.xml) of Chinese Wikipedia to generate a graph of all Wikipedia concepts, we call it pre-stage. Then we enter stage one by selecting three top-level nodes as roots, traversing the graph generated in the pre-stage using BFS-like algorithm to form spanning trees and computing rough domain relatedness of these nodes at the same time. Finally, in stage two we use the novel Modified Explicit Semantic Analysis method combined with the results we got in stage one to compute the ultimate domain relatedness. The experimental results shows that our method could get a high-quality domain-specific lexical repository.
Nowadays, rapid evolution of computers and mobile devices has caused the explosive increase in network traffic. So it becomes more and more necessary to archive network traffic for analyzing network events and a lot of emerging applications. Compression is fundamental for traffic archival solution to save the storage space, and indexing is effective to accelerate search queries for archive of traffic data. In this paper, we propose BreadZip (blocks row-reordering and adaptive index zip), a combination of initial traffic data and index compression. BreadZip has three main advantages. 1) to improve compressing efficiency and reduce memory footprint, traffic data is reordered in sequence and divided into fixed-size blocks; 2) to accelerate queries, an improved bitmap indexes with smaller volume than traditional will be introduced; 3) to save space, both traffic blocks and bitmap indexes are compressed in different simple run-length encoding methods respectively. Finally, our empirical results on network traffic from CAIDA (Cooperative Association for Internet Data Analysis) show that our solution can significantly reduce the volume of traffic data, while simultaneously preserving the ability to perform selectively queries with response times in seconds.
Semantic relatedness measures are used in many applications in natural language processing and we propose a Wikipedia-based method to compute it. Unlike existed methods that only focus on a small section of Wikipedia (e.g. info box or hyperlinks), our method makes full use of the rich information contained in the Wikipedia page and could get a higher accuracy within reasonable time. In our method, we first use some special sections (e.g. synonyms and hyponyms) in the Wikipedia page to judge whether two concepts are closely related. If they are not, we then use pattern matching to find whether they are related through usual relatedness (e.g. "a part of", "result in", and "is a member of"). And if the relatedness score hasn't been computed out through former steps, we then use a method which makes some improvement on the famous explicit semantic analysis method to compute the relatedness.
Core metadata for seafloor observatory network is the important basis for seafloor observatory data directory, which is supposed to run in the undersea observatory network system. It is also an important way to make data of ocean shared in the marine information construction of China. The paper analyzes the content of related metadata and standards at home and abroad. Combined with the features of the seafloor observatory data, based on defining principles and requirements of core metadata for seafloor observatory network, content of core metadata for seafloor observatory is determined. This paper provides the core metadata structure design using the Universal Modeling language (UML). The detail definitions of core metadata elements are given in data dictionary.
Domain-specific corpus can be used to build domain ontology, which is used in many areas such as IR, NLP and web Mining. We propose a multi-root based method to build a domain-specific corpus making use of Wikipedia resources. First we select some top-level nodes (Wikipedia category articles) as root nodes and traverse the Wikipedia using BFS-like algorithm. After the traverse, we get a directed Wikipedia graph (Wiki-graph). Then an algorithm mainly based on Kosaraju Algorithm is proposed to remove the cycles in the Wiki-graph. Finally, topological sort algorithm is used to traverse the Wiki-graph, and ranking and filtering is done during the process. When computing a node’s ranking score, the in-degree of itself and the out-degree of its parents are both considered. The experimental evaluation shows that our method could get a high-quality domain-specific corpus
0引言随着计算机学科与各学科之间的关系越来越紧密,信息技术对人类社会全方位的渗透,计算机公共课程的基础性地位不言而喻,但由于缺乏类似数学和物理课程那样在教育体系中连贯的课程设置,很多高校计算机公共课程的设置长期以来一直处于20世纪90年代初的科普培训模式框架中。部分任课教师由于缺乏对科研环节的投入,很难将不断发展变化的信息技术很好地在一门基础课中进行诠释,在某种程度上浪费了基础课量大面广的教学资源。
随着面向服务思想在业务领域的不断应用,服务的数量急剧增加,海量业务服务形成了一个庞大复杂的业务网络,服务间的动态组合和协作形成了面向服务的业务生态系统,为了高效地进行服务选择,充分发挥服务关联对于服务选择和组合的重要作用,提出了一种支持服务应用关联的服务选择方法。基于服务选择历史信息构建服务关联网络;基于服务关联网络,建立支持服务应用关联的服务选择目标规划模型,并采用遗传算法对所建模型进行求解。通过实验进行了算例分析,验证了所提方法的可行性和有效性。
Analytical customer relationship management (CRM) systems, which can discover knowledge from huge amount of data, play a crucial role in decision support. However, while most researches in CRM mainly focus on data mining, relatively few research papers on detailed architectures for analytical CRM systems have been published. In this paper, a practical process driven architecture of building analytical CRM systems is proposed. This architecture is based on the typical CRM analysis process and integrates with multilevel secure (MLS) data model to ensure the efficiency and timeliness of provision of right information to the right users. Furthermore, through a case study of a real-world analytical CRM implementation project in the bank industry, an analytical CRM solution using the proposed architecture has been developed and received approvable ratings from our client, which demonstrates the feasibility of this architecture.
Tsinghua University currently offers a general education curriculum on computer fundamentals for non-computer majors. This curriculum plays an important role in helping students establishing a complete knowledge structure in computer technology and in developing students' interdisciplinary research capabilities. In the survey of needs for such a curriculum at Tsinghua university, eleven major areas' special characteristics and their latest needs for new talents were carefully analyzed. A new curriculum was designed based on the results from the survey. This new curriculum features a set of solid foundational courses, a clear organization of concepts by types and by levels, applications orientated. After two years of implementation, the new curriculum has been shown to be very effective.
《网页设计与制作》是当前热门课程,该课程集知识和技能于一体,实践性很强。我们以一个学习任务为主线贯穿整个学习活动,精心设计教学内容和与之配套的实验教学,采取在机房上课、边讲边练的方式和案例教学法,锻炼了学生解决实际问题的能力、自主学习能力和协作学习能力。
Customer churn analysis and prediction play an important role in customer relationship management and improve benefit of enterprise.A Support Vector Machine model is established to predict customer churn.Customer churn characteristic is presented in this paper.According to the churn data which is large scale and imbalance,this paper presents a two-class model based on improved SVM to predict customer churn.The class weighted SVM model CW-SVM is presented,and the accuracy is improved by adjusting the class weight and the position of boundary.The efficiency is improved by translating the SVM to the Core Vector Machine and a new algorithms CWC-SVM is presented.The arithmetic performance is better than others based on the test of real credit debt data set in the commercial bank.
提出了国家精品课程优质资源共享系统的推荐系统框架。在介绍了整个框架的应用背景之后,逐步地将推荐系统的概念,应用情况以及目前的研究进展加以总结性介绍。对于组成本推荐系统框架的两个模块分别作了概要的原理性说明,在说明中,重点就系统选型,模块选择的根据进行了详细论述。最后通过简单而具有说明力的例子来帮助理解,同时说明了此框架的可行性并简单描述了框架的实现情况。