A 45-year-old woman complained dry cough 2 months ago and had felt chest discomfort, fever, and abdominal pain for half a month. The fever was recurrent with the highest temperature of 39°C. She had no history of long-term alcohol consumption but underwent thymectomy due to thymic tumor and myasthenia gravis 9 years ago. Pyridostigmine bromide has been stopped for 3 years. On admission, she presented anemic appearance, mild epigastric tenderness without rebound tenderness. Laboratory test showed that serum amylase and lipase were much higher than three folds upper limitation of normal. Leucocyte count 2.71 × 10/L with normal classification, hemoglobin 91 g/L, platelet count 88 × 10/L, and D-dimer 13.78 mg/L FEU. Calcium and triglycerides were within normal ranges. Fecal occult blood was positive. Contrast-enhanced computed tomography (CECT) revealed her pancreas as sausage-like and is obviously swollen (Fig. 1a). The density of the pancreatic parenchyma was uneven with large volume of reduced enhancement (Fig. 1b–e). Due to severe necrosis of the pancreatic body, her pancreas was almost split into several parts (Fig. 1e). The severe pancreatic necrosis did not cause massive exudation as usual because the subacute necrotic lesion was wrapped within capsule-like rim (Fig. 1a–e). Moderate pericardial effusion and mild pleural effusion on the left side were detected (Fig. 1f). The tests for her IgG4, tuberculosis, and autoimmune disease were negative. All of cultures for blood, sputum, and urine were negative as well. Bone marrow biopsy showed normal morphology of the three lineage cells. What kind of pancreatitis is this? Is there any relationship between the pancreatitis and pericardial effusion? The sausage-like pancreas in CECT looks a little bit as IgG4-related pancreatitis. But it does not result in pancreatic necrosis normally. Gastroscopy found an ulcer in her gastric body (Fig. 2a). The ulcer biopsy presented active gastritis with polypoid hyperplasia. The brown granules in the gastric mucosa of immunohistochemistry staining indicated positive for Epstein–Barr virus (EBV, Fig. 2b,c). Her serum level of EBV-DNA was high as 4.03 × 10 copies/mL. These evidences suggest that the patient suffered from systemic infection of EBV by which pancreatitis and pericarditis became the parts of this disease. Although supportive treatments were conducted, the patient gradually developed multiple organ failure and died 20 days later. This EBV infection unusually occurred in an adult case possibly due to decreased immunity caused by thymus surgery. There have been less than 20 cases of EBV-pancreatitis reported in last 50 years’ literature. Although pancreas is rare to be involved in systemic infection of EBV, EBV-pancreatitis may be developed into severe necrotic one but with less exudation and subacute onset. The prognosis of such necrotic pancreatitis is poor because of less understanding of its pathophysiology and no effective treatment against to EBV.
A 68‐year‐old female farmer has suffered persistent moderate dull pain in the right lower abdomen for 19 months, without any other suggestive symptoms and signs. Laboratory data and imaging did not show evidence of infectious, autoimmune, or neoplastic diseases. Colonoscopy showed a normal terminal ileum but scattered erythematous mucosa and purulent secretions in the ileocecum (Fig. 1a). Biopsy revealed chronic inflammation with granulomas (Fig. 1b). Immunohistochemical staining of EBV‐encoded RNA, acid‐fast staining, and polymerase chain reaction test for tuberculosis were negative. Crohn’s disease (CD) was considered, and mesalazine was given for 10 months. However, her symptom was not improved, but an egg‐sized mass appeared at her right lumbar region. Abdominal computed tomography and intestinal contrast ultrasonography suspected the presence of ileocecal fistula with sinus formation extending to lumbar abdominal wall subcutaneously (Fig. 1c,d). Percutaneous puncture of the lumbar sinus was performed, but culture of the extracted pus was negative. Cefotaxime, ceftezole, and amoxicillin sulbactam were applied with standard dose for 7 days consecutively as empirical treatment, which turned out ineffective. Repeated colonoscopy showed nodular lesions in the ileocecum, and the biopsy revealed chronic inflammation. Crohn’s disease diagnosis for this patient was doubtful due to her old age of onset with striking intestinal fistula but mild symptoms, atypical findings of colonoscopy. Then ileocecal resection was indicated for ileocecal fistula with sinus formation. Surgery found that congestive and swelled ileocecum adhered closely to posterior and lateral abdominal wall. There was a fistula from ileocecum to lumbar back skin, with 1.5 cm in width and 5 cm in depth. Postoperative pathology revealed the presence of flaky mycelioid structures and granulomas (Fig. 2a). Methenamine silver staining confirmed the presence of actinomycetes (Fig. 2b). After surgery, combined with intravenous therapy of benzylpenicillin sodium (5.6 million IU × q6h for 30 days), the patient recovered completely. Her lumbar sinus was also healed. During 1‐year follow‐up, there were no signs of recurrence. The diagnosis of actinomycosis was finally confirmed. Actinomycosis, as an opportunistic infection, often affected patients with immunodeficiency, cancer, trauma, or malnutrition. However, it has also been reported in some healthy individuals. Abdominal actinomycosis is an unusual form of visceral involvement and has been reported rarely. As a group of anaerobic Gram‐positive bacteria normally colonizing the gastrointestinal tract, actinomycetes may be pathogenic when mucosal injury occurs. Appendix, cecum, and colon are the most commonly affected sites in the digestive tract. It is usually difficult to diagnose actinomycosis because of the lower positive rate of actinomycetes with routine cultivation method. Methenamine silver staining of tissue section by which typical characteristics of actinomycetes can be detected may be helpful to the suspected cases. Moreover, actinomycete infection is characterized by significant fibrosis, formation of granulomas, and transmural and suppurative inflammation with abscess and fistula, which may mimic CD. However, CD is always presented in young patients with typical findings of colonoscopy. In rare cases, actinomycete infection may also be a comorbidity in patients with CD, making the disease more complicated. In conclusion, abdominal actinomycosis may mimic fistulizing CD, and it should be considered in the differential diagnosis of CD.
The results of Human Genome Project promote the development of bioinformatics.Searching disease genes that have function correlations,also called similar phenotype genes,based on the strategy of disease phenome similarity becomes an emerging research topic due to its important research value and wide range of applications.However,in biomedical field,there is no previous work that applies computer methods to search similar phenotype genes via a network consists of "gene-disease-phenotype" relations.To fill the gap,in this study,a disease information network containing three heterogeneous nodes (i.e.,gene,disease,and phenotype) is built by making use of a disease open database.In addition,an algorithm,called gSim-Miner,is designed for the search of similar phenotype genes via the disease information network.Pruning strategies based on the characteristics of disease phenotype data are proposed to improve the efficiency of gSim-Miner.Experiments on real-world data sets demonstrate that the disease information network is feasible,and gSim-Miner is effective,efficient and extensible.
Internet网络大数据与日俱增,当前亟需设计出能够处理大规模半结构化和无结构化文本数据的新型聚类方法.现有工作的不足体现在:应用的文本集较为单一,对半结构和无结构的Web文本进行聚类的准确性较低,当文档规模较大时聚类的时效性无法得到保证.针对上述不足,提出新的基于群体智能的文本聚类模型Switch(a Swarm intelligence based text clustering algorithm),支持包括藏文、汉文、英文等多语言的文本聚类.基本思想为:构建文本的向量空间模型,借助自然语言处理和数据预处理技术得到由特征向量构成的文本集合;对群体智能文本聚类算法的参数进行初始化,不同智能体可以在二维文本空间上任意移动,计算其所在网格区域文本与其他样本的相似度,利用概率转换函数求取智能体拿起和放下样本的概率,进而实现文本聚类.提出分布式动态文本流聚类的multi-agent架构,将这一架构应用于群体智能文本聚类算法中,分布式工作环境被设计成相互通信的软agents集合,设计了相似度计算,智能体状态感知,文本解析三类智能体.通过解决智能体状态同步、处理器负载均衡和处理器之间通信的代价问题,将计算任务分成不同子任务,在多处理器上分布执行.此外,阐述了基于multi-agent的分布式群体智能文本聚类方法的工作原理,给出一种分布式通信架构,各种智能体相互通信,相互协作完成文本聚类工作.基于multi-agent通过JADE(Java Agent Development Framework)中间件实现集群上的分布式文本聚类,优势在于:分布式计算和大内存处理较单机具有更好的处理能力,借助JADE中间件能够使智能体间相互通信及协作,实现高效的文本聚类.在大量真实的半结构化包含藏文、汉文和英文多语言的Web文本数据集上进行实验,以藏文为例,实验结果表明:相比于k-means和单节点上的群体智能聚类算法,提出的分布式架构下文本聚类算法准确性平均高出12.2%和3.8%,时间代价平均缩减了73.0%和50.6%.在n个节点集群下agents数量介于150~250之间时,文本聚类时间代价近似可以达到单节点的1/n.
Calculation of the information network data cube (InfoNetCube) is the foundation of information online analytical processing.However,different from the traditional data cube,InfoNetCube consists of multiple lattices in which each cuboid contains a topic graph (or graph measurement),thus the storage consumption overhead is two orders of magnitude more than that of traditional data cube.How to materialize the specified cuboids or lattice rapidly and efficiently in the information network is a quite challenging research issue.In this paper,a novel InfoNetCube materializing strategy for information network is proposed based on dialysis computing.By leveraging the anti-monotonicity of topic graph measurement in the information and topology dimensions,a dialysis based space pruning algorithm is constructed to rapidly dialysis out the hidden sub graph,cuboids and lattices.Experimental results show that the proposed partial materialization algorithm outperform the cube based partial materialization strategy,saving almost 75% aggregation time.
Dynamic information network is a new challenging problem in the field of current complex networks.The evolution of dynamic networks is temporal,complex and changeable.Structure is the basic characteristics of the network,and is also the basis of network modeling and analysis.The study of the network structure evolution is of great importance in getting a comprehensive understanding of the behavior trend of complex systems.This paper introduces "role" to quantify the structure of dynamic network and proposes a role-based model.To predict the role distributions of dynamic network nodes in future time,the presented framework views role prediction as a multi-target regression problem,extracts properties from historical snapshot sub-network,and predicts the future role distributions of dynamic network nodes.The paper then proposes a multi-target regression based role prediction (MTR-RP) method for dynamic network.This method not only overcomes the drawback of the existing methods which operate on transfer matrix while ignoring the time factor,but also takes into account of possible dependencies between multiple forecast targets.Experiments results show that MTR-RP has better and more stable prediction capability compared with the existing methods.
As the size of networks grows larger, traditional community discovery algorithms cannot effectively and efficiently process the large-scale network data.Based on the Spark distributed graph computing model, this study proposes a parallel algorithm for discovering communities in large-scale complex networks, called DBCS(Discovering Big Community on Spark).The proposed approach employs the basic idea of clustering method beyond modularity, which first calculates the increment of the modularity between the node pairs, and then iteratively finds the maximum modularity increment among all the node pairs.Lastly, it merges the node pairs, and updates the modularity increment of the remaining nodes, in order to identify the communities in large-scale complex networks.Extensive experiments are conducted on several real and synthetic network datasets and the results demonstrate that DBCS can effectively deal with the problem of partitioning the large-scale networks that does not make sense for traditional algorithms.In particular, it only takes about four minutes to handle more than one million nodes for community discovery.In addition, the time cost is reduced to 1/20 of the parallel algorithm based on Hadoop.The accuracy is improved by 7.4% when compared to traditional community discovery algorithms.
Anomaly detection,which is used in a variety of applications,attracts attention both in industry and academia.Among numerous methods for anomaly detection,the Isolation Forest algorithm,whose characteristics include high efficiency,sound detection accuracy,has wide real-world applications.However,the conventional Isolation forest algorithm can hardly deal with large-scale data sets.To break this limitation,we propose a cloud computing platform based algorithm.Specifically,we design and implement a parallel algorithm for anomaly detection based on Isolation Forest,named PIFH,using the Hadoop distributed storage system and the MapReduce distributed computational framework.By parallelizing the processes of detection model construction and anomaly evaluation,its efficiency is improved,and the application range is also extended.Experiments using real-world data sets demonstrate that the proposed algorithm is efficient and scalable.
Background: Anterolateral thigh flap is perfect for reconstructing maxillofacial soft tissue defects. This flap has been widely used by clinicians, but often causes operation difficulties because of vascular variation. Thus sometimes anteromedial thigh was used as new donor site temporarily when the vascular anatomic variation of anterolateral thigh perforator flap induced a failure in the flap harvest. Objectives: To discuss the anatomy and application of anteromedial thigh flap in 13 cases. Methods: We collect 13 cases of the anteromedial thigh flap transplantation during 2009 to 2015. Seven of them were elected due to the error of the ultrasonic test, three of them were elected because of the diameter being too small, other three cases were elected due to the failure of iatrogenic behaviours. Findings and Conclusions: Seven of the cases had vessels directly to the skin, six of them were intramuscular perforators. 10 of the cases were bilobate flaps, and three of them were folding flaps. All of the 13 cases survived with no vascular crisis occurring. The follow-up results after three to six months were satisfactory.
DSP (distinguishing sequential pattern) is a kind of sequence such that it occurs frequently in the sequence set of target class, while infrequently in the sequence set of non-target class.Since distinguishing sequential patterns can describe the differences between two sets of sequences, mining of distinguishing sequential patterns has wide applications, such as building sequence classifiers, characterizing biological features of DNA sequences, and behavior analysis for specified group of people.Compared with mining distinguishing sequential patterns satisfying the predefined support thresholds, mining distinguishing sequential patterns with top-k contrast measure can avoid setting improper support thresholds by users.Thus, it is more user-friendly.However, the conventional algorithm for mining top-k DSPs cannot deal with the sequence data set with large-scale.To break this limitation, a parallel mining method using Spark, named SP-kDSP-Miner (Spark based top-k DSP miner), is designed for mining top-k distinguishing sequential patterns from large-scale sequence data set.Furthermore, in order to improve the efficiency of SP-kDSP-Miner, a novel candidate pattern generation strategy and several pruning strategies, as well as a parallel computing method for the contrast scores of candidate patterns are proposed considering the characteristics of Spark structure.Experiments on both real-world and synthetic data sets demonstrate that SP-kDSP-Miner is effective, efficient and scalable.
Sequential data are prevalent in many real world applications. The quality evaluation on sequential data, which attracts the attentions from both academic research and industry fields, is important and prerequisite for extracting knowledge from the sequential data. Recently, a method using the probabilistic suffix tree has been proposed for evaluating the sequential data quality. However, this method cannot deal with the large-scale data set. To break this limitation, this paper proposes a Spark-based algorithm, called STALK (sequential data quality evaluation with Spark), for evaluating the quality of large-scale sequential data. Moreover, this paper uses the novel pruning strategies to improve the efficiency of STALK. Specifically, on the Spark platform, the large-scale sequential data are efficiently used to generate model, and the data quality of query sequence can be evaluated according to the generated model rapidly. Experiments on real-world sequential data sets demonstrate that STALK is effective, efficient and scalable.
Contrast patterns describe differences between two or more data sets or data classes; they have been proven to be useful for solving many kinds of problems, such as building accurate classifiers, defining clustering quality measures, and analyzing disease subtypes. This article investigates the mining of a new kind of contrast patterns, namely discriminating inter‐attribute functions (DIFs), which represent arithmetic‐expression‐based inter‐attribute relationships that distinguish classes of data. DIFs are an expressive and practical alternative of item‐based contrast patterns and can express discriminating relationships such as “ weight /( height ) 2 is more likely to be ≤25 in one class than in another class.” Besides introducing the DIF mining problem, this article makes theoretical and algorithmic contributions on the problem. We prove that DIF mining is MAX SNP‐hard. Regarding how to efficiently mine DIFs, we present a set of rules to prune the search space of arithmetic expressions by eliminating redundant ones (equivalent to some others). We give two algorithms: one for finding all DIFs satisfying given thresholds and another for finding certain optimal DIFs using genetic computation techniques. The former is useful when the number of attributes is small, whereas the latter is useful when that number is large; both use the redundant arithmetic‐expression pruning rules. A performance study shows that our techniques are effective and efficient for finding DIFs.
Distinguishing sequential patterns with gap constraints are very useful for identifying important features for discriminating one class of sequences from sequences of other classes. A gap constraint should be predefined by users using the proposed mining methods for distinguishing sequential patterns.It is difficult for users to set suitable gap constraints without enough priori knowledge.As a result,useful patterns may be missed.To deal with this problem, this paper presents an algorithm for mining distinguishing sequential patterns with compact gap constraints.The proposed algorithm runs well without a predefined gap constraint.Instead,it computes the most suitable gap constraint for each candidate pattern.In addition,three pruning rules are designed to improve the efficiency of the algorithm.Experiments on protein sequences, DNA sequences,and activity sequences confirmed the effectiveness and efficiency of the proposed algorithm.
We tackle the novel problem of mining contrast subspaces. Given a set of multidimensional objects in two classes \(C_+\) and \(C_-\) and a query object \(o\), we want to find the top-\(k\) subspaces that maximize the ratio of likelihood of \(o\) in \(C_+\) against that in \(C_-\). Such subspaces are very useful for characterizing an object and explaining how it differs between two classes. We demonstrate that this problem has important applications, and, at the same time, is very challenging, being MAX SNP-hard. We present CSMiner, a mining method that uses kernel density estimation in conjunction with various pruning techniques. We experimentally investigate the performance of CSMiner on a range of data sets, evaluating its efficiency, effectiveness, and stability and demonstrating it is substantially faster than a baseline method.
Distinguishing sequential pattern (DSP) mining has been widely employed in many applications, such as building classifiers and comparing/analyzing protein families. However, in previous studies on DSP mining, the gap constraints are very rigid – they are identical for all discovered patterns and at all positions in the discovered patterns, in addition to being predetermined. This paper considers a more flexible way to handle gap constraint, allowing the gap constraints between different pairs of adjacent elements in a pattern to be different and allowing different patterns to use different gap constraints. The associated DSPs will be called DSPs with flexible gap constraints. After discussing the importance of specifying/determining gap constraints flexibly in DSP mining, we present GepDSP, a heuristic mining method based on Gene Expression Programming, for mining DSPs with flexible gap constraints. Our empirical study on real-world data sets demonstrates that GepDSP is effective and efficient, and DSPs with flexible gap constraints are more effective in capturing discriminating sequential patterns.
Mining contrast sequential patterns, which are sequential patterns that characterize a given sequence class and distinguish that class from another given sequence class, has a wide range of applications including medical informatics, computational finance and consumer behavior analysis. In previous studies on contrast sequential pattern mining, each element in a sequence is a single item or symbol. This paper considers a more general case where each element in a sequence is a set of items. The associated contrast sequential patterns will be called itemset-based distinguishing sequential patterns (itemset-DSP). After discussing the challenges on mining itemset-DSP, we present iDSP-Miner, a mining method with various pruning techniques, for mining itemset-DSPs that satisfy given support and gap constraint. In this study, we also propose a concise border-like representation (with exclusive bounds) for sets of similar itemset-DSPs and use that representation to improve efficiency of our proposed algorithm. Our empirical study using both real data and synthetic data demonstrates that iDSP-Miner is effective and efficient.
Community structure is an important feature that exists extensively in real-world complex networks. Tradi-tional community evolution studies are limited to the analysis on single-level communities, and have some defects, such as the evolutionary regularities revealing and algorithms stability, etc. To handle the problems, this paper pro-poses an information networks community trend prediction method based on structure analysis. The method obtains community hierarchies by hierarchical clustering, matches communities with different structures in adjacent network snapshots, therefore relatively overcomes the difficulty of overlooking the influence of sudden outside events, and provides possibility for the structure based community evolution analysis. The method is applied in two real-world datasets, and the experimental results show that the work in this paper greatly improves the algorithm efficiency and stability.
For intelligent transportation systems, digital military battlefield and driver assistance systems, it is of great practical value to predict the trajectories of moving objects with uncertainty in a real-time, accurate and reliable fashion. Intelligent trajectory prediction can not only provide accurate location-based services, but also monitor and estimate traffic to suggest the best path, and as such becomes an active research direction. Aiming to overcome the drawbacks of the existing methods, a new trajectory prediction model based on Gaussian mixture models called GMTP is proposed. The new model contains the following essential phases: (1) modeling the complex motion patterns based on Gaussian mixture models, (2) calculating the probability distribution of different types of motion patterns by using Gaussian mixture model in order to partition trajectory data into distinct components, and (3) inferring the most possible trajectories of moving objects via Gaussian process regression. The GMTP algorithm is naturally a Gaussian nonlinear statistical probability model and the advantage of the proposed model is that the result is not only a predicted value, but also a whole distribution beyond the future trajectories, therefore making it possible to infer the location in regard to some motion patterns, e.g., uniformly accelerated motion, by using statistical probability distribution. Extensive experiments are conducted on real trajectory data sets and the results show that the prediction accuracy of the GMTP algorithm is improved by 22.2% and 23.8%, and the runtime can be reduced by 92.7% and 95.9% on average, respectively, when compared to the Gaussian process regression model and Kalman filter prediction algorithm with similar parameter setting.
随着电子商务的发展,许多购物网站都提供商品评论作为用户购物的决策参考。由于商品评论具有海量、冗余、不规范的特点,用户难以在短时间内浏览所有商品评论,更难以基于评论内容发现商品对比特征。对此,设计了top-k显露模式挖掘算法,并将此算法应用于商品评论对比分析,实现了用户购物决策支持系统——Review Scope。Review Scope能够从不同商品的评论中发现特定商品的对比评论,并以此作为购物决策可视化地提供给用户。基于京东商城真实商品评论数据的实验结果表明Review Scope具有有效、灵活、用户友好的特点。
Limin Xiang合作论文数Kyushu Sangyo University Department of Information Science
Teシスsocial intelligence disciplines Home6