The objective of the study was to investigate the impact of Covid-19 on mental health and specifically depression for the Swiss population. Data from Swiss Household Panel surveys from 2019 (no Covid) with 8841 individuals and from 2020 (Covid) with 15882 individuals about their life and health conditions, in addition to data from statistics on Income and living Conditions (SILC) conducted by Swiss Federal Statistical Office (FSO) from 2019 (no covid) and 2020 (covid) with approx. 18000 individuals along with Covid data from Swiss Surveillance Systems (CH-SUR) from 2020 were pre-processed, cleaned, transformed and finally integrated to a unique relational database. Apriori-algorithm for association analysis was applied and led to finding some interesting rules. The result of this study are 3 rules extracted from the database showing the correlation between Covid-19 and depression in the year of 2020 in Switzerland and compared with depression data in the year 2019 (before Covid-19). The found rules show that the Swiss population's life and health condition was not significantly impacted during the first year of Covid-19 in 2020. In our future work, we intend to use different data sources like social media instead of surveys to verify the gained results in this study.
This article demonstrates that using data mining methods such as Weighted Association Rule Mining (WARM) on an integrated Swiss database derived from a Swiss national dietary survey (menuCH) and 25 years of Swiss demographical and health data is a powerful way to determine whether a specific population subgroup is at particular risk for developing a lifestyle disease based on its food consumption patterns. The objective of the study was to discover critical food consumption patterns linked with lifestyle diseases known to be strongly tied with food consumption. Food consumption databases from a Swiss national survey menuCH were gathered along with data of large surveys of demographics and health data collected over 25 years from Swiss population conducted by Swiss Federal Office of Public Health (FOPH). These databases were integrated and reported in a previous study as a single integrated database. A data mining method such as WARM was applied to this integrated database. A set of promising rules and their corresponding interpretation was generated. As an example, the found rules of the sample show that the consumption of alcohol in small quantities does not have a negative impact on health, whereas the consumption of vegetables is important for the supply of vitamins of the B group, which help the energy metabolism to pro-vide energy. These vitamins are particularly lacking in alcoholics and should then be taken with supplements. Another finding is that dietary supplements do little specially by diabetes. Applying WARM algorithm was beneficial for this study since no interesting rules were pruned out early and the significance of the rules could be highly increased as compared to a previous study using pure Apriori Algorithm.
Nowadays, all kinds of service-based organizations open online feedback possibilities for customers to share their opinion. Swiss National Railways (SBB) uses Facebook to collect commuters' feedback and opinions. These customer feedbacks are highly valuable to make public transportation option more robust and gain trust of the customer. The objective of this study was to find interesting association rules about SBB's commuters pain points. We extracted the publicly available FB visitor comments and applied manual text mining by building categories and subcategories on the extracted data. We then applied Apriori algorithm and built multiple frequent item sets satisfying the minsup criteria. Interesting association rules were found. These rules have shown that late trains during rush hours, deleted but not replaced connections on the timetable due to SBB's timetable optimization, inflexibility of fines due to unsuccessful ticket purchase, led to highly customer discontent. Additionally, a considerable amount of dis-satisfaction was related to the policy of SBB during the initial lockdown of the Covid-19 pandemic. Commuters were often complaining about lack of efficient and effective measurements from SBB when other passengers were not following Covid-19 rules like public distancing and were not wearing protective masks. Such rules are extremely useful for SBB to better adjust its service and to be better prepared by future pandemics.
Objective: The objective of the study was to link Swiss food consumption data with demographic data and 30 years of Swiss health data and apply data mining to discover critical food consumption patterns linked with 4 selected chronical diseases like alcohol abuse, blood pressure, cholesterol, and diabetes. Design: Food consumption databases from a Swiss national survey menu CH were gathered along with data of large surveys of demographics and health data collected over 30 years from Swiss population conducted by Swiss Federal Office of Public Health (FOPH). These databases were integrated and Frequent Pattern Growth (FP-Growth) for the association rule mining was applied to the integrated database. Results: This study applied data mining algorithm FP-Growth for association rule analysis. 36 association rules for the 4 investigated chronic diseases were found. Conclusions: FP-Growth was successfully applied to gain promising rules showing food consumption patterns lined with lifestyle diseases and people’s demographics such as gender, age group and Body Mass Index (BMI). The rules show that men over 50 years consume more alcohol than women and are more at risk of high blood pressure consequently. Cholesterol and type 2 diabetes is found frequently in people older than 50 years with an unhealthy lifestyle like no exercise, no consumption of vegetables and hot meals and eating irregularly daily. The intake of supplementary food seems not to affect these 4 investigated chronic diseases.
Objective: The objective of the study was to integrate a large database from Swiss nutrition national survey (menu-CH) with 5 extensive databases derived from 5 consecutive Swiss health national surveys from 1992 to 2012 for data mining purposes. Each database has additionally a demographic base data. An integrated Swiss database is built to later discover critical food consumption patterns linked with lifestyle diseases known to be strongly tied with food consumption and compare the derived rules with the rules resulted with a previous study which used a significantly smaller database. Design: Swiss nutrition national survey (menuCH) with approx. 2000 respondents from two different surveys, one by Phone and the other by questionnaire along with Swiss health national surveys from 1992 to 2012 with over than 100000 respondents were preprocessed, cleaned, transformed and finally integrated to a unique relational database. Results: The result of this study is an integrated relational database from the Swiss nutritional and 20 years of Swiss health data.
Objective: The objective of the study was to integrate two big databases from Swiss nutrition national survey (menuCH) and Swiss health national survey 2012 for data mining purposes. Each database has a demographic base data. An integrated Swiss database is built to later discover critical food consumption patterns linked with lifestyle diseases known to be strongly tied with food consumption. Design: Swiss nutrition national survey (menuCH) with approx. 2000 respondents from two different surveys, one by Phone and the other by questionnaire along with Swiss health national survey 2012 with 21500 respondents were pre-processed, cleaned and finally integrated to a unique relational database. Results: The result of this study is an integrated relational database from the Swiss nutritional and health databases. Keywords—Health informatics, data mining, nutritional and health databases, nutritional and chronical databases.
Background: This article demonstrates that using data mining methods such as association analysis on an integrated Swiss database derived from a Swiss national dietary survey (menuCH) and Swiss demographical and health data is a powerful way to determine whether a specific population subgroup is at particular risk for developing a lifestyle disease based on its food consumption patterns. Objective: The objective of the study was to use an integrated database of dietary and health data from a large group of Swiss population to discover critical food consumption patterns linked with lifestyle diseases known to be strongly tied with food consumption. Design: Food consumption databases from a Swiss national survey menuCH were gathered along with corresponding large survey of demographics and health data from Swiss population conducted by Swiss Federal Office of Public Health (FOPH). These databases were integrated and reported in a previous study as a single integrated database. A data mining method such as A-priori association analysis was applied to this integrated database. Results: Association mining analysis was used to incorporate rules about food consumption and lifestyle diseases. A set of promising preliminary rules and their corresponding interpretation was generated, which is reported in this paper. As an example, the found rules of the sample show that smoking is relatively irrelevant to the high blood pressure and Diabetes, whereas consuming vegetables at regular basis reduces the risk of high Cholesterol. Conclusions: Association rule mining was successfully used to describe and predict rules linking food consumption patterns with lifestyle diseases. The gained association rules reveal that the appearance of the mutually independent nutritional characteristics in the rules are equally distributed. Furthermore, most of the sample show no chronical diseases as they smoke little and exercise regularly, which can be interpreted that sport is a strong preventive factor for chronical/lifestyle diseases. Nevertheless, a small percentage of the sample shows chronic illnesses due to unhealthy eating. Further research should consider the weighting of chronic diseases' characteristics for them not to be pruned out early by data mining computation.
In this paper, we present a robust algorithm to character classes. In addition, our experiments have shown recognize extracted text from grocery product images captured and image processing. HE use of digital cameras to capture text from natural Sciences,3005 Bern, Switzerland (e-mail: farshideh.einsele@bfh.ch). Science, University of Central Florida, Orlando, FL 32816, USA (e-mail: foroosh@cs.ucf.edu). World Academy of Science, Engineering and Technology International Journal of Computer, Electrical, Automation, Control and Information Engineering Vol:9, No:1, 2015 159 International Scholarly and Scientific Research & Innovation 9(1) 2015 scholar.waset.org/1999.4/10000257 In te rn at io na l S ci en ce I nd ex , C om pu te r an d In fo rm at io n E ng in ee ri ng V ol :9 , N o: 1, 2 01 5 w as et .o rg /P ub lic at io n/ 10 00 02 57 into Fourier domain and ignore the phase information to gain rotation invariant features. Scale and translation invariance is gained with normalization and considering the centroid as the center of the coordinate system. The experimental results show that the centroid and complex coordinate signatures have a high precision and recall rate whereas the curvature and cumulative angular functions deliver unreliable results. Dionisio et al. in [11] also report a contour-based shape classification technique based on polygon approximation that is invariant under rotation and scaling. The vertices of polygon approximation are formed by high curvature points of the profile and are selected by the Fourier transform of the object contour. A series of features are computed from the polygonal approximation and a minimum distance classifier is used for object recognition. Although such contour-based invariants deliver promising results using Fourier descriptors for character recognition and are also reported in the survey of Trier et al. [2], the reported test and training databases are synthetically deformed patterns and the features are invariant with respect to translation, scaling, rotation and do not consider other transformations coming from real world captured images (e.g. shearing, shadowing, bad illumination and perspective distortion). Besides Trier et al.state in [2] that a statistical classification system should consider the so called curse of dimensionality meaning that it should be training-based containing a minimum number of patterns that is 8-10 times bigger than the number of the chosen features. As already stated, when dealing with a database containing real world character images, database generation is an expensive and time-consuming drawback. To sum up, the above mentioned works have been performed either by considering a synthetically degraded database or are training-based approaches. In this paper, we present a method for camera-based character recognition that uses a small real world database extracted from images of grocery products captured by a cellular phone with a resolution of 5 mega pixels. The presented method does not need a training set that should rely on the previously described term of curse of dimensionality. Therefore our proposed method can be applied directly to the extracted text with no use of cost-intensive image enhancement algorithms and delivers promising results. The remainder of this paper is organized as follows: in section II, we introduce the specificities of text in product images. Section III reports about our proposed character recognition algorithm including used feature extraction and classification methods. Section IV presents our evaluation results and section V is about our conclusions and a short sketch of our future works. II. CHALLENGES OF PRODUCT TEXT RECOGNITION We extract text from images taken with cell phones from grocery products. Text extraction from camera-based images is a relatively well researched area with plenty of existing works in the literature [12]. However, text extraction methods from camera-based images is tightly related to a specific application and there does not exist a valid generic method for the extraction of text within different camera-based scenarios. We therefore use a text extraction algorithm that has been developed for the specificities of the text from grocery products and is explained in detail in [13]. The resolution of the used cell phone camera is 5 mega pixels and the images are taken from different angles with the camera having a similar distance to products as the one a common grocery shopper would have when he crosses grocery aisles. The extracted text has mostly a height between 20-50 pixels and characters can be mostly labeled and segmented using connected components algorithms. Table I shows some extracted words in our database.
In this paper, we present a robust algorithm to character classes. In addition, our experiments have shown recognize extracted text from grocery product images captured and image processing. HE use of digital cameras to capture text from natural Sciences,3005 Bern, Switzerland (e-mail: farshideh.einsele@bfh.ch). Science, University of Central Florida, Orlando, FL 32816, USA (e-mail: foroosh@cs.ucf.edu). World Academy of Science, Engineering and Technology International Journal of Computer, Information, Systems and Control Engineering Vol:9 No:1, 2015 159 International Scholarly and Scientific Research & Innovation 9(1) 2015 In te rn at io na l S ci en ce I nd ex V ol :9 , N o: 1, 2 01 5 w as et .o rg /P ub lic at io n/ 10 00 02 57 into Fourier domain and ignore the phase information to gain rotation invariant features. Scale and translation invariance is gained with normalization and considering the centroid as the center of the coordinate system. The experimental results show that the centroid and complex coordinate signatures have a high precision and recall rate whereas the curvature and cumulative angular functions deliver unreliable results. Dionisio et al. in [11] also report a contour-based shape classification technique based on polygon approximation that is invariant under rotation and scaling. The vertices of polygon approximation are formed by high curvature points of the profile and are selected by the Fourier transform of the object contour. A series of features are computed from the polygonal approximation and a minimum distance classifier is used for object recognition. Although such contour-based invariants deliver promising results using Fourier descriptors for character recognition and are also reported in the survey of Trier et al. [2], the reported test and training databases are synthetically deformed patterns and the features are invariant with respect to translation, scaling, rotation and do not consider other transformations coming from real world captured images (e.g. shearing, shadowing, bad illumination and perspective distortion). Besides Trier et al.state in [2] that a statistical classification system should consider the so called curse of dimensionality meaning that it should be training-based containing a minimum number of patterns that is 8-10 times bigger than the number of the chosen features. As already stated, when dealing with a database containing real world character images, database generation is an expensive and time-consuming drawback. To sum up, the above mentioned works have been performed either by considering a synthetically degraded database or are training-based approaches. In this paper, we present a method for camera-based character recognition that uses a small real world database extracted from images of grocery products captured by a cellular phone with a resolution of 5 mega pixels. The presented method does not need a training set that should rely on the previously described term of curse of dimensionality. Therefore our proposed method can be applied directly to the extracted text with no use of cost-intensive image enhancement algorithms and delivers promising results. The remainder of this paper is organized as follows: in section II, we introduce the specificities of text in product images. Section III reports about our proposed character recognition algorithm including used feature extraction and classification methods. Section IV presents our evaluation results and section V is about our conclusions and a short sketch of our future works. II. CHALLENGES OF PRODUCT TEXT RECOGNITION We extract text from images taken with cell phones from grocery products. Text extraction from camera-based images is a relatively well researched area with plenty of existing works in the literature [12]. However, text extraction methods from camera-based images is tightly related to a specific application and there does not exist a valid generic method for the extraction of text within different camera-based scenarios. We therefore use a text extraction algorithm that has been developed for the specificities of the text from grocery products and is explained in detail in [13]. The resolution of the used cell phone camera is 5 mega pixels and the images are taken from different angles with the camera having a similar distance to products as the one a common grocery shopper would have when he crosses grocery aisles. The extracted text has mostly a height between 20-50 pixels and characters can be mostly labeled and segmented using connected components algorithms. Table I shows some extracted words in our database.
In this paper, we present a robust algorithm to character classes. In addition, our experiments have shown recognize extracted text from grocery product images captured and image processing. HE use of digital cameras to capture text from natural Sciences,3005 Bern, Switzerland (e-mail: farshideh.einsele@bfh.ch). Science, University of Central Florida, Orlando, FL 32816, USA (e-mail: foroosh@cs.ucf.edu). World Academy of Science, Engineering and Technology International Journal of Computer, Information, Systems and Control Engineering Vol:9 No:1, 2015 159 International Scholarly and Scientific Research & Innovation 9(1) 2015 In te rn at io na l S ci en ce I nd ex V ol :9 , N o: 1, 2 01 5 w as et .o rg /P ub lic at io n/ 10 00 02 57 into Fourier domain and ignore the phase information to gain rotation invariant features. Scale and translation invariance is gained with normalization and considering the centroid as the center of the coordinate system. The experimental results show that the centroid and complex coordinate signatures have a high precision and recall rate whereas the curvature and cumulative angular functions deliver unreliable results. Dionisio et al. in [11] also report a contour-based shape classification technique based on polygon approximation that is invariant under rotation and scaling. The vertices of polygon approximation are formed by high curvature points of the profile and are selected by the Fourier transform of the object contour. A series of features are computed from the polygonal approximation and a minimum distance classifier is used for object recognition. Although such contour-based invariants deliver promising results using Fourier descriptors for character recognition and are also reported in the survey of Trier et al. [2], the reported test and training databases are synthetically deformed patterns and the features are invariant with respect to translation, scaling, rotation and do not consider other transformations coming from real world captured images (e.g. shearing, shadowing, bad illumination and perspective distortion). Besides Trier et al.state in [2] that a statistical classification system should consider the so called curse of dimensionality meaning that it should be training-based containing a minimum number of patterns that is 8-10 times bigger than the number of the chosen features. As already stated, when dealing with a database containing real world character images, database generation is an expensive and time-consuming drawback. To sum up, the above mentioned works have been performed either by considering a synthetically degraded database or are training-based approaches. In this paper, we present a method for camera-based character recognition that uses a small real world database extracted from images of grocery products captured by a cellular phone with a resolution of 5 mega pixels. The presented method does not need a training set that should rely on the previously described term of curse of dimensionality. Therefore our proposed method can be applied directly to the extracted text with no use of cost-intensive image enhancement algorithms and delivers promising results. The remainder of this paper is organized as follows: in section II, we introduce the specificities of text in product images. Section III reports about our proposed character recognition algorithm including used feature extraction and classification methods. Section IV presents our evaluation results and section V is about our conclusions and a short sketch of our future works. II. CHALLENGES OF PRODUCT TEXT RECOGNITION We extract text from images taken with cell phones from grocery products. Text extraction from camera-based images is a relatively well researched area with plenty of existing works in the literature [12]. However, text extraction methods from camera-based images is tightly related to a specific application and there does not exist a valid generic method for the extraction of text within different camera-based scenarios. We therefore use a text extraction algorithm that has been developed for the specificities of the text from grocery products and is explained in detail in [13]. The resolution of the used cell phone camera is 5 mega pixels and the images are taken from different angles with the camera having a similar distance to products as the one a common grocery shopper would have when he crosses grocery aisles. The extracted text has mostly a height between 20-50 pixels and characters can be mostly labeled and segmented using connected components algorithms. Table I shows some extracted words in our database.
In this paper, we introduce and evaluate a system capable of recognizing ultra low resolution words extracted from images such as those frequently embedded on web pages. The design of the system has been driven by the following constraints. First, the system has to recognize small font sizes where antialiasing and resampling procedures have been applied. Such procedures add noise on the patterns and complicate any a priori segmentation of the characters. Second, the system has to be able to recognize any words in an open vocabulary setting, potentially mixing different languages. Finally, the training procedure must be automatic, i.e. without requesting to extract, segment and label manually a large set of data. These constraints led us to an architecture based on ergodic HMMs where states are associated to the characters. We also introduce several improvements of the performance increasing the order of the emission probability estimators and including minimum and maximum duration constraints on the character models. The proposed system is evaluated on different font sizes and families, showing good robustness for sizes down to 6 points.
We report in this paper on significant improvements that we have been including in our HMM-based system for recognition of ultra low resolution, antialiased text with small point sizes such as those frequently found in web images. First we are proposing a fully automatic training procedure where no a priori knowledge of font metrics is needed. This means that our system can be potentially built to recognize any font. Second, the system’s performance can be boosted by using mixtures of Gaussians to model the probability density functions of the HMMs. Third, we show that these improvements allow the system to handle large vocabulary size up to 60’000 wordswith few degradation of the accuracy. We report on these results for 2 different font families, namely a serif and a sans serif font. We also report on different HMMs topologies and conclude on the benefits of using minimum duration to model the characters composing the words.
In this paper, we present a HMM based system that is used to recognize ultra low resolution text such as those frequently embedded in images available on the web. We propose a system that takes specifically the challenges of recognizing text in ultra low resolution images into account. In addition to this, we show in this paper that word models can be advantageously built connecting together sub-HMM-character models and inter-character state. Finally we report on the promising performance of the system using HMM topologies which have been improved to take into account the presupposed minimum length of each character.
Current Web indexing technologies suffer from a severe drawback due to the fact that Web documents often present textual information that is encapsulated in digital images and therefore not available as actual coded text. Moreover such images are not suited to be processed by existing OCR software, since they are generally designed for recognizing binary document images produced by scanners with resolutions between 200-600 dpi, whereas text embedded in web images is often anti-aliased and has generally a resolution between 72 and 90 dpi. The presented paper describes two preliminary studies about character identification at very low resolution (72 dpi) and small font sizes (3-12 pts). The proposed character identification system delivers identification rates up to 99.93% for 12psila600 isolated character samples and up to 99.89% for 300psila000 character samples in context.
Current OCR technology does not allow to accurately recognizing small text images, such as those found in web images. Our goal is to investigate new approaches to recognize very low resolution text images containing anti-aliased character shapes.This paper presents a preliminary study on the variability of such characters and the feasibility to discriminate them by using geometrical features. In a first stage we analyze the distribution of these features. In a second stage we present a study on the discriminative power for recognizing isolated characters, using various rendering methods and font properties. Finally we present interesting results of our evaluation tests leading to our conclusion and future focus.
Jean Hennebert合作论文数;Software Engineering Unit
Business Information System Institute
University of Applied Science - HES-SO ; Wallis4