From a cross-linguistic perspective, language models are interesting because they can be used as idealised language learners that learn to produce and process language by being trained on a corpus of linguistic input. In this paper, we train different language models, from simple statistical models to advanced neural networks, on a database of 41 multilingual text collections comprising a wide variety of text types, which together include nearly 3 billion words across more than 6,500 documents in over 2,000 languages. We use the trained models to estimate entropy rates, a complexity measure derived from information theory. To compare entropy rates across both models and languages, we develop a quantitative approach that combines machine learning with semiparametric spatial filtering methods to account for both language- and document-specific characteristics, as well as phylogenetic and geographical language relationships. We first establish that entropy rate distributions are highly consistent across different language models, suggesting that the choice of model may have minimal impact on cross-linguistic investigations. On the basis of a much broader range of language models than in previous studies, we confirm results showing systematic differences in entropy rates, i.e. text complexity, across languages. These results challenge the long-held notion that all languages are equally complex. We then show that higher entropy rate tends to co-occur with shorter text length, and argue that this inverse relationship between complexity and length implies a compensatory mechanism whereby increased complexity is offset by increased efficiency. Finally, we introduce a multi-model multilevel inference approach to show that this complexity-efficiency trade-off is partly influenced by the social environment in which languages are used: languages spoken by larger communities tend to have higher entropy rates while using fewer symbols to encode messages.
One of the fundamental questions about human language is whether all languages are equally complex. Here, we approach this question from an information-theoretic perspective. We present a large scale quantitative cross-linguistic analysis of written language by training a language model on more than 6500 different documents as represented in 41 multilingual text collections consisting of ~ 3.5 billion words or ~ 9.0 billion characters and covering 2069 different languages that are spoken as a native language by more than 90% of the world population. We statistically infer the entropy of each language model as an index of what we call average prediction complexity. We compare complexity rankings across corpora and show that a language that tends to be more complex than another language in one corpus also tends to be more complex in another corpus. In addition, we show that speaker population size predicts entropy. We argue that both results constitute evidence against the equi-complexity hypothesis from an information-theoretic perspective.
Languages employ different strategies to transmit structural and grammatical information. While, for example, grammatical dependency relationships in sentences are mainly conveyed by the ordering of the words for languages like Mandarin Chinese, or Vietnamese, the word ordering is much less restricted for languages such as Inupiatun or Quechua, as these languages (also) use the internal structure of words (e.g. inflectional morphology) to mark grammatical relationships in a sentence. Based on a quantitative analysis of more than 1,500 unique translations of different books of the Bible in almost 1,200 different languages that are spoken as a native language by approximately 6 billion people (more than 80% of the world population), we present large-scale evidence for a statistical trade-off between the amount of information conveyed by the ordering of words and the amount of information conveyed by internal word structure: languages that rely more strongly on word order information tend to rely less on word structure information and vice versa. Or put differently, if less information is carried within the word, more information has to be spread among words in order to communicate successfully. In addition, we find that-despite differences in the way information is expressed-there is also evidence for a trade-off between different books of the biblical canon that recurs with little variation across languages: the more informative the word order of the book, the less informative its word structure and vice versa. We argue that this might suggest that, on the one hand, languages encode information in very different (but efficient) ways. On the other hand, content-related and stylistic features are statistically encoded in very similar ways.
The wdlpOst dictionary writing system to be presented in this paper has been developed for the specific purposes of a lexicographical project on German loanwords in the East Slavic languages Russian, Belarusian, and Ukrainian. The project's main objectives are (i) to document those loanwords for which a cognate lexical borrowing from German is known in Polish and (ii) to establish possible borrowing pathways for these lexical items. In the first phase of the project, the collaborative client/server architecture of the wdlpOst system has been used for excerpting detailed lexicographical information from a large range of historical and contemporary East Slavic dictionaries, taking the entries in a large dictionary of German loanwords in Polish as a common frame of reference. For the project's second phase, the wdlpOst system provides innovative tooling for compiling entries of the East Slavic loanwords. Most importantly, the numerous word sense definitions for a set of cognate loanwords, as excerpted from different lexicographical sources, are mapped onto a system of newly defined cross-language word senses; in a similar vein, the phonemic and graphemic variation in the loanwords and their derivatives is captured through a tool that abstracts from dictionary-specific idiosyncrasies.
and structural abnormalities in the Rasmussen Center for Cardiovascular Disease Prevention The 9 non-invasive tests, in addition to BP, include small artery and large artery elasticity, exercise BP response, optic fundus photography, carotid ultrasound, microalbuminuria, ECG, left ventricular ultrasound and NT-ProBNP measurement. A Disease Score (DS) of 0-18 is based on abnormalities in these 9 tests, with DS of 0-2 associated with no future morbid events, DS of 3-5 with morbid events delayed by at least 4 years and DS of 6 experiencing a high morbid event rate beginning immediately. Results: After excluding 632 individuals taking anti-hypertensive drugs 1184 were classified by BP: normal (<120/80 mmHg), n1⁄4403; pre-hypertensive stage I (BP range, 120-129/80-84), n1⁄4476); prehypertensive stage II, (BP range, 130-139/85-89), n1⁄4185); hypertensive (BP 140/ 90), n1⁄4120). Mean Disease Score in these 4 groups were 2.5 2.2, 3.5 2.5, 4.3 2.8 and 5.6 3.3, respectively, thus confirming the population relationship between BP and disease. However, DS of 6 was present in over 10% of the normal group, 22% of prehypertension stage I, 30% of prehypertension stage II and 50% of hypertension. The distribution of BP in each of the Disease Score groups was largely overlapping. Conclusion: BP is not a reliable guide to individual risk for cardiovascular morbid events. Early non-invasive disease assessment provides the opportunity for individualized management of early disease regardless of BP.
In this paper, the authors use the 2012 log files of two German online dictionaries (Digital Dictionary of the German Language and the German Version of Wiktionary) and the 100,000 most frequent words in the Mannheim German Reference Corpus from 2009 to answer the question of whether dictionary users really do look up frequent words, first asked by de Schryver et al. (2006). By using an approach to the comparison of log files and corpus data which is completely different from that of the aforementioned authors, we provide empirical evidence that indicates - contrary to the results of de Schryver et al. and Verlinde/Binon (2010) - that the corpus frequency of a word can indeed be an important factor in determining what online dictionary users look up. Finally, we incorporate word class Information readily available in Wiktionary into our analysis to improve our results considerably.
We start by trying to answer a question that has already been asked by de Schryver et al. (2006): Do dictionary users (frequently) look up words that are frequent in a corpus. Contrary to their results, our results that are based on the analysis of log files from two different online dictionaries indicate that users indeed look up frequent words frequently. When combining frequency information from the Mannheim German Reference Corpus and information about the number of visits in the Digital Dictionary of the German Language as well as the German language edition of Wiktionary, a clear connection between corpus and look-up frequencies can be observed. In a follow-up study, we show that another important factor for the look-up frequency of a word is its temporal social relevance. To make this effect visible, we propose a de-trending method where we control both frequency effects and overall look-up trends.
CI 1⁄4 confidence interval; CV 1⁄4 cardiovascular; CHD 1⁄4 coronary heart disease; MACE 1⁄4 major adverse CV events (stroke, CHD, CV death; estimated for some trials); HF 1⁄4 heart failure. These data suggest that, in clinical trials, amlodipine significantly increased the risk of heart failure, but prevented other cardiovascular endpoints at least as well as other regimens, including those that started with a diuretic.
The cellular composition of the tumor microenvironment may affect survival in diffuse large B-cell lymphoma (DLBCL). We performed immunostains for 2 stromal cell markers, CD68 and SPARC (secreted protein, acidic and rich in cysteine), in 262 patients with DLBCL treated with rituximab and cyclophosphamide, doxorubicin, vincristine, and prednisone (CHOP) or CHOP-like therapies. Patients with any SPARC+ cells in the microenvironment had a significantly longer overall survival, and patients with high SPARC positivity in the microenvironment also had a significantly longer event-free survival. Survival differences were mainly due to the prognostic effect of SPARC+ cells in activated B-cell (ABC)-type DLBCL, with no effect found in the germinal center B-cell-type DLBCL. Of clinical features examined, only the number of extranodal sites was significantly associated with SPARC expression. Multivariate analysis revealed that SPARC expression predicted patient survival independent of the International Prognostic Index or tumor cell of origin. SPARC expression in the microenvironment of DLBCL can be used for prognostic purposes, determining a subgroup of patients with ABC DLBCL who have significantly longer survival. More aggressive chemotherapy protocols should be considered for patients with ABC DLBCL without SPARC+ stromal cells. CD68 expression by cells in the microenvironment did not predict survival.
Objective: To investigate the safety and effectiveness of 2 simple discharge regimens for use in patients with type 2 diabetes mellitus (DM2) and severe hyperglycemia, who present to the emergency department (ED) and do not need to be admitted.Methods: We conducted an 8-week, open-label, randomized controlled trial in 77 adult patients with DM2 and blood glucose levels of 300 to 700 mg/dL seen in a public hospital ED. Patients were randomly assigned to receive glipizide XL, 10 mg orally daily (G group), versus glipizide XL, 10 mg orally daily, plus insulin glargine, 10 U daily (G+G group). The primary outcome was to maintain safe fasting glucose and random glucose levels of <350 and <500 mg/dL up to 4 weeks and <300 and <400 mg/dL, respectively, thereafter and to have no return ED visits (responders).Results: Baseline characteristics were similar between the 2 treatment groups. The primary outcome was achieved in 87% of patients in both treatment groups. The enrollment mean blood glucose values of 440 and 467 mg/dL in the G and G+G groups, respectively, declined by the end 0 of week 1 to 298 and 289 mg/dL and by week 8 to 140 and 135 mg/dL, respectively. Homeostasis model assessment of beta-cell function and early insulin response improved 7-fold and 4-fold, respectively, in responders at the end of the 8-week study.Conclusion: Sulfonylurea with and without use of a small dose of insulin glargine rapidly improved blood glucose levels and beta-cell function in patients with DM2. Use of sulfonylurea alone once daily can be considered a safe discharge regimen for such patients and an effective bridge between ED intervention and subsequent follow-up. (Endocr Pract. 2009;15:696-704)
Background-Measurement of carotid intima-media thickness (CIMT) has been validated as a measure of atherosclerosis and as a predictor of future cardiovascular events. Compared with glimepiride, pioglitazone has been shown to slow the progression of atherosclerosis measured by CIMT in patients with type 2 diabetes mellitus.Methods and Results-We evaluated individual cardiovascular risk factors as predictors of the change in CIMT produced by pioglitazone treatment by determining whether their addition to a baseline model resulted in loss of significance for the treatment effect on CIMT. Pioglitazone treatment led to improvement in levels of multiple cardiovascular risk markers, including high-sensitivity C-reactive protein, apolipoprotein B, apolipoprotein A1, high-density lipoprotein (HDL) cholesterol, triglyceride, insulin, and free fatty acid. At 24 weeks, there were significant differences in HDL cholesterol, triglyceride, total cholesterol, low-density lipoprotein cholesterol, insulin, body mass index, hip circumference, and high-sensitivity C-reactive protein between the pioglitazone and glimepiride treatment groups. After adjustment for 24-week on-treatment values of cardiovascular risk factors, only inclusion of the changes in HDL cholesterol and insulin significantly impacted the magnitude and significance of the treatment effect on CIMT. Furthermore, irrespective of treatment assignment, increased HDL cholesterol at 24 weeks was a significant predictor of reduced CIMT progression at 72 weeks.Conclusions-The beneficial effect of pioglitazone on HDL cholesterol at 24 weeks predicted its beneficial effect for reducing CIMT progression at 72 weeks. Changes in HDL cholesterol at 24 weeks, irrespective of treatment, predicted less progression of CIMT at 72 weeks. These results suggest that suppression of atherosclerosis with pioglitazone therapy is linked to its ability to raise HDL cholesterol.
We evaluated correlates of coronary atherosclerosis, measured by coronary artery calcium, in a racially diverse group of male and female subjects with type 2 diabetes. Age, systolic blood pressure, sex, and race/ethnicity were significant determinants of coronary artery calcium. Among lipoproteins, cholesterol level contained in a particle excluded from direct measures of LDL and HDL cholesterol (designated triglyceride-rich lipoprotein cholesterol) was most strongly linked to coronary artery calcium. Neither inflammatory markers nor metabolic factors correlated with coronary artery calcium in models adjusted for age and sex, but measures of adipose distribution did. Waist-to-hip ratio and the ratio of visceral to total abdominal tissue were positively associated with coronary artery calcium. In fully adjusted multivariate models, the relationship of adiposity measures to coronary artery calcium was no longer significant after inclusion of apolipoprotein B or triglyceride-rich lipoprotein cholesterol. Traditional risk factors and race/ethnicity remain important correlates of coronary artery calcium in a cohort at elevated risk of cardiovascular disease because of type 2 diabetes. Adiposity measures are significantly associated with coronary artery calcium score, but their importance may be largely explained by apolipoprotein B or triglyceride-rich lipoprotein cholesterol.
CONTEXT:Carotid artery intima-media thickness (CIMT) is a marker of coronary atherosclerosis and independently predicts cardiovascular events, which are increased in type 2 diabetes mellitus (DM). While studies of relatively short duration have suggested that thiazolidinediones such as pioglitazone might reduce progression of CIMT in persons with diabetes, the results of longer studies have been less clear.OBJECTIVE:To evaluate the effect of pioglitazone vs glimepiride on changes in CIMT of the common carotid artery in patients with type 2 DM.DESIGN, SETTING, AND PARTICIPANTS:Randomized, double-blind, comparator-controlled, multicenter trial in patients with type 2 DM conducted at 28 clinical sites in the multiracial/ethnic Chicago metropolitan area between October 2003 and May 2006. The treatment period was 72 weeks (1-week follow-up). CIMT images were captured by a single ultrasonographer at 1 center and read by a single treatment-blinded reader using automated edge-detection technology. Participants were 462 adults (mean age, 60 [SD, 8.1] years; mean body mass index, 32 [SD, 5.1]) with type 2 DM (mean duration, 7.7 [SD, 7.2] years; mean glycosylated hemoglobin [HbA1c] value, 7.4% [SD, 1.0%]), either newly diagnosed or currently treated with diet and exercise, sulfonylurea, metformin, insulin, or a combination thereof.INTERVENTIONS:Pioglitazone hydrochloride (15-45 mg/d) or glimepiride (1-4 mg/d) as an active comparator.MAIN OUTCOME MEASURE:Absolute change from baseline to final visit in mean posterior-wall CIMT of the left and right common carotid arteries.RESULTS:Mean change in CIMT was less with pioglitazone vs glimepiride at all time points (weeks 24, 48, 72). At week 72, the primary end point of progression of mean CIMT was less with pioglitazone vs glimepiride (-0.001 mm vs +0.012 mm, respectively; difference, -0.013 mm; 95% confidence interval, -0.024 to -0.002; P = .02). Pioglitazone also slowed progression of maximum CIMT compared with glimepiride (0.002 mm vs 0.026 mm, respectively, at 72 weeks; difference, -0.024 mm; 95% confidence interval, -0.042 to -0.006; P = .008). The beneficial effect of pioglitazone on mean CIMT was similar across prespecified subgroups based on age, sex, systolic blood pressure, duration of DM, body mass index, HbA(1c) value, and statin use.CONCLUSION:Over an 18-month treatment period in patients with type 2 DM, pioglitazone slowed progression of CIMT compared with glimepiride.TRIAL REGISTRATION:clinicaltrials.gov Identifier: NCT00225264