Radio frequency identification (RFID) technology is an automatic identification method, which relies on the use of radio repeaters called RFID tags to quickly store and retrieve data. Because RFID tags do not need linear contact when communicating with readers, it is possible to collect a large amount of data in a short time. However, the collected data also produce problems such as false negative readings, false positive readings, duplicated readings, out-of-order readings and so on. In this case, how to efficiently clean the large-scale RFID data in short time has become an important research topic in the field of database. This paper mainly summarizes the existing RFID data cleaning technology. Firstly, the relevant definitions and descriptions of RFID system and RFID data cleaning problem are given, and typical datasets and evaluation criteria are listed. Then the existing RFID data cleaning work is compared and summarized in detail from the classification, subcategory, basic idea, advantages, limitations, application scenarios and other aspects of related technologies. At the same time, the relevant application systems are compared and analyzed. Then, for the key problems of the false negative reading, false positive reading, duplicated reading and out-of-order reading, the existing studies are compared and summarized in detail. Finally, this paper proposes next five research directions worthy of paying attention to in the field of RFID data cleaning, such as the construction of RFID original data and benchmark dataset, the cleaning strategy of encryption and privacy protection data, the accuracy of data collection, the timeliness of cleaning results, and scene self-learning.
Coronary heart disease is the first killer of human health. At present, the common way of coronary heart disease diagnosis is coronary angiography. This method is a kind of surgery that could cause some physical damage to the patient, and it could cause some complications and adverse reactions. Furthermore, coronary angiography is expensive. However, the heart color Doppler echocardiography report, blood biochemical indicators and other basic information can reflect the degree of heart damage in patients. Therefore, this paper proposes a combined reinforcement multitask progressive networks (CRMPN) model to predict the grade of coronary heart disease through heart color Doppler echocardiography report, blood biochemical indicators and ten basic body information items about the patient. In this model, the first step is performing deep reinforcement learning (DRL) pre-training through asynchronous advantage actor-critic (A3C). Training data is adopted to optimize the recurrent neural networks (RNN) that parameterizes the stochastic policy. In the second step, we use soft parameter sharing module, hard parameter sharing module and progressive deep network to predict coronary heart disease. The experimental results show that after DRL pre-training, the multiple tasks in the model interact with each other and learn together to achieve satisfactory results and outperform other state-of-the-art methods.
数据开放为海量数据中的价值得以最大化利用提供了可能,然而数据的安全性问题却成开放共享的最大阻碍。针对开放数据集查询分析结果包含的隐私信息,基于其与数据访问行为的直接联系,设计一种安全规则存储结构。提出面向自然语言的安全需求描述接口,对待保护的隐私信息进行灵活、方便的描述,进一步提出自然语言隐私需求描述到安全规则的自动转换方法。在此基础上建立数据安全自动审核模型,该模型根据数据拥有者的安全需求对访问者的数据行为进行审核,实现在数据隐私不泄露的前提下保证数据开放的最大化。通过真实数据集上的实验表明:安全规则能够准确地捕获数据提供者的隐私保护需求,并能够有效地保证数据的安全性。
SummaryEntity recognition plays an important role in building the electronic medical records (EMRs) based medical knowledge graph, which is significant for building Clinical decision support (CDS) system. Cross‐disease clinical documents are context‐related and have different interrelated semantic structures, which bring challenges for entity recognition using traditional methods. In order to solve these problems, this paper proposes a co‐training based entity recognition approach for cross‐disease clinical documents. In this model, we first build partial annotation corpus of the single disease using dependency syntax analysis and the medical statement rule unifies. Then, according to the partial annotation corpus of different diseases, the sentence level features are extracted through the Bi‐LSTM layer with memory unit and CRF methods, which optimize the whole sequence and improve the combination probability of sequence labels. Finally, the results with higher confidence are selected by cross feedback to label the corpus, which enlarges the size of corpus and improves the accuracy of the document entity recognition. The experiment result proves the availability and high efficiency of our method.
结合解决我国高血压疾病管理存在的主要问题,应用现代信息技术,开发基于互联网运行的纵向协同管理系统,即可尝试构建一套从下而上的自我教育、预防干预、健康体检、医院治疗和从上而下的医院治疗、社区随访、家庭帮助、自我管理的纵向协同管理体系,降低高血压疾病的患病率、高增长趋势和高危害性;提高高血压管理的知晓率、治疗率和控制率,改善患者生活质量.
String search is an important branch of pattern matching for information retrieval in various fields. In the past four decades, the research importance has been attached on skipping more unnecessary characters to improve the search performance, and never taken into consideration on large scale of data. In this paper, two major achievements are contributed. At first, we propose a Quick Search algorithm for data Stream (QSS) on a single machine to support string search in a large text file, as opposed to previous researches that limits to a bound memory. For the next, we implement the search algorithm on MapReduce framework to improve the velocity of retrieving the search results. The experiments demonstrate that our approach is fast and effective for large files.
Asystemic solution for the application construction of self-governing openness of data based on the new data management pattern of self-governing openness of data was proposed.The whole framework,system functions and interfaces of the solution were provided.Then the key points about logical data organization,security requirements description and query requirements description technologies were designed and investigated for the data provider and data consumer,respectively.Finally,self-governing openness of data practices in the field of medical was discussed,and the framework of tiered diagnosis and treatment system application by using the proposed technologies was given.
Multi-center clinical research is the main approach for multi-center,multi-disciplinary,to develop some collaborative clinical researches on the same clinical issues.The traditional multi-center clinical research mainly has the disadvantages that small sample size and the clinical research is relatively closed and the degree of openness is not high.Therefore,the newly emerging technology,such as big data and cloud computing was combined to integrate clinical centers of physically dispersed hospitals into a logical and unified clinical data.On this basis,a multi-center clinical big data application platform was constructed.First,the overall framework of the multi-center clinical big data platform was designed,and then the subsystems of the platform were elaborated in detail.Finally,the deep application of clinical big data platform was introduced.
The characters of big data are volume,variety,value,velocity,and common hardware and open source.Aiming at the system inefficiency and limited scalability of traditional relational database in big data analysis,this paper presen-ted an algorithm of Hash joins in MapReduce distributed environment based on column-store by introducing MapReduce computing model.First of all,this paper proposed the design of large data-oriented distributed computing models.Then, it proposed the partition aggregation and the heuristic optimization strategy to realize the implementation of Hash join algorithm.Lastly,the experiments evaluated execution time and load capacity.The results show that the proposed method is effective and can provid good scalability in big data analysis.
完善财务绩效管理评价标准,提高财务信息化管理水平是相关企业和单位追求的目标.文章以我国的三甲医院为主要实证对象,构建了财务信息化绩效物元模型,采用可拓评价方法建立了三甲医院财务信息化绩效评价指标体系结构,对我国83个三甲医院的财务信息化绩效水平进行科学而有效的实证研究.实证结果表明,提出的财务信息化绩效物元模型及其可拓评价方法能够准确反映出我国三甲医院财务信息化绩效现状.
As a new method to describe entities and their relationships,knowledge graph has been paid more and more attention in the medical field.However,most of the knowledge of the medical knowledge graph is derived from the open medical literature,and less related to the EMR electronic medical records.EMR electronic medical records cover the whole process of patient diagnosis and treatment with a wealth of medical facts,which is an important source of knowledge of medical knowledge graph.Therefore,this paper takes the specific disease of breast tumor as an example.According to the basic principle of knowledge graph technology,we firstly gave the definition of knowledge of breast tumors.Combined with the actual EMR electronic medical records data set of Ruijin Hospital Affiliated to Shanghai Jiaotong University School of Medicine,the knowledge of breast cancer medical facts was extracted from EMR by means of knowledge extraction technology.On this basis,a method for constructing knowledge map of breast tumors is proposed.
Hierarchical medical system is the concerned medical system today, which is of great significance for reducing the pressure on the patient and tertiary institutions. A large number of ordinary patients go to the experts in tertiary institutions because they don't trust the basic-level hospital and worry about misdiagnose, which results in shortage of medical resources in tertiary institutions and waste of medical resources in the basic medical institution. In this paper, we design a thyroid disease evaluation model based on deep learning for blood tests and thyroid ultrasound report by analyzing the clinical data of thyroid patients in a top three comprehensive hospital. The model can assess the severity of the thyroid disease and give referral advice which can help doctor in the basic medical institution to make decision, so as to improve the accuracy of the doctor's judgment and guide patients to basic medical institution. This is a new way to promote the implementation of hierarchical medical system.
Extracting clinical entities and their relations from clinical texts is a preliminary task for constructing medical knowledge graph. Existing end-to-end models for extracting entity and relation have limited performance in clinical text because they rarely take both latent syntactic information and the effect of context information into account. Thus, this paper proposed a context-aware end-toend neural model for extracting relations between entities from clinical texts with two level attention. We show that entity-level attention effectively acquires more syntactic information by learning a weighted sum of child word nodes rooted at the target entities. Meanwhile, sub sentence-level attention in an effort to capture the interactions between the target entity pairs and context entity pairs by assigning weights of each context representation within one sentence. Experiments on real-world clinical texts from Shanghai Ruijin Hospital demonstrate that our model significantly gains better performance in the application of clinical texts compared with existing joint models.
访问控制技术已广泛应用于数据库安全领域,但是它无法防范SQL注入、内部人员权限滥用等非法行为。针对这些问题,提出可信计算环境下的数据库强制行为控制(MBC)模型,判断用户提交事务的可信性;设计并实现了可信数据库控制基(TDCB)原型对行为策略进行完整性度量。实验结果显示,MBC能够阻止不可信的事务执行,有效解决内部人员运行非法事务的问题;TDCB能够检测出MBP被后门用户非法篡改,并禁止执行相应的事务以免造成损失。
The optimal path planning for robot patrol is a combinatorial optimization problem that aims to solve the smallest Hamiltonian circle of a complete graph. Such problems are typical NP-hard problems, and the computational complexity of the existing precise algorithm increases exponentially with the expansion of the problem scale. Even if a considerable amount of running time is employed, it is difficult to obtain the global optimal solution. This paper presents a robot patrol path planning method based on combined deep reinforcement learning (DRL). This method combines reinforcement learning and neural networks to perform path optimization. The first step is performing DRL pre-training through asynchronous advantage actor-critic (A3C). Training data is adopted to optimize the recurrent neural networks (RNN) that parameterizes the stochastic policy. The second step is effective planning (EP) without pre-training. Using the expected reward objective, it iteratively optimizes the RNN parameters on test instances. The experimental results show that combined deep reinforcement learning method can jump out of the local optimal solution and outperform other state-of-the-art methods.
基于新经济地理学理论,考察空间集聚对企业全要素生产率的影响关系及内在机制.采用1998-2007年《中国工业企业数据库》及对应的《中国城市统计年鉴》构建面板数据模型进行计量检验.研究结果显示:空间集聚程度与企业全要素生产率之间存在典型的“倒U”型关系;空间集聚有利于通过增强企业竞争行为提升资源配置效率来促进企业全要素生产率.认为,当前中国大多数城市并未呈现过度集聚现象,城市化进程中应该实现市场在资源配置中的决定性作用.通过实现要素在空间层面的自由流动,在加强空间集聚外部性的同时,促进资源配置效率,从而实现微观企业生产效率的提升.
Punching of textile uppers is an important process in the shoe production line, which determines the accuracy of positioning and affects all subsequent shoemaking steps. At present, there are few researches on the 3-DOF punching robot in this field. The critical value of the punching force of the end mechanism is not clear, and the punching path planning is not adopted with a suitable algorithm, resulting in large power consumption and low work efficiency. Therefore, this paper describes the design of a punching robot to meet the action requirements of punching the textile uppers and complete the path planning to improve efficiency. The first step is to model a 3-DOF robot for textile upper punching including rail type and articulated type to meet the needs of punching movements. The second step is to gain the lowest critical value of the pressure with ANSYS by simulating deformation process when the punch punches in the vamp until breaks it. And this outcome is the theory basis for motor selection and robot design. The third step is multi-point punching path planning, where a combined deep reinforcement learning (DRL) method is proposed. The DRL pre-training is performed through asynchronous advantage actor -critic (A3C). Training data is adopted to optimize the recurrent neural networks (RNN) that parameterizes the stochastic policy. Then the inspire planning (IP) algorithm is conducted to get the optimal path. The experimental results show that this method can jump out of the local optimal solution and outperform other state-of-the-art methods.
Entity recognition of clinical documents is a primary task to extract information from unstructured clinical documents. Traditional entity recognition methods extract entities in a supervised learning framework which needs a large scale of labeled corpus as the training samples. However, clinical documents in real world are unlabeled. To construct a large scale of labeled corpus by manual is time-consuming. Semi-supervised learning that relies on small-scale corpus can solve such problem. Thus, this paper proposes an entity recognition model of clinical documents based on self-training framework. In such framework, we first establish partial annotation corpus through the way of dependency syntax analysis and the medical statement rule unifies. Then, a hybrid model of CNN-LSTM-CRF is proposed to label the unlabeled data in an end-to-end way. Specially, we will use CNN to embed characters in clinical document and use Bi-LSTM to extract the sentence-level feature. At the moment, we use CRF remedies the shortage of LSTM which further combined with the combination probability of CRF and the advantages of optimizing the whole sequence. Finally, the results of entity recognition with higher confidence level are fed back by self-training to expend size of corpus which improves the accuracy of the document entity recognition. The experiment result proves the availability and high efficiency of this model.
Wireless sensor networks (WSNs) have gained huge popularity in various new fields. The direct use of each of these sensors individually is to detect its surrounding conditions such as temperature, pressure, sound, motion etc. whereas a collection of such nodes finds application in various large scale management systems such as healthcare, disaster and traffic management programs. In most of these systems, the sensors are located at points which are not physically protected and can hence fall prey to several security threats and attacks very easily. The limited memory, power and capacity of such sensors make it more difficult to introduce advanced, heavy-weight algorithms for securing them against such attacks. Further, the criticality of the applications using WSN makes such security threats more dangerous. In this paper, we show the results of combining a novel detection algorithm with the results of a unique resolution technique of the sinkhole attack when the network is routed using Collection Tree Protocol (CTP).