
Forage plants are considered one of the main factors for the development of livestock worldwide, for presenting high potential for phytomass production, drought tolerance, high energy value, large water reserve and easy propagation. Forage cactus stands out for its tolerance to water deficit. Aimed to evaluate the initial performance of the morphometric characteristics of Giant Sweet clone (N. cochenillifera) submitted to water and saline stresses using response surface analysis. Design used was completely randomized in a 4x4 factorial scheme, composed of four levels of water replacement, based on crop evapotranspiration (ETc): (25%.ETc, 50%.ETc, 75%.ETc and 100%.ETc) and four levels of water salinity (0, 2, 4 and 8 dS/m), with four repetitions totalling 64 experimental units. The following morphometric characteristics were evaluated: plant height, length, width, thickness, number of cladodes and area of cladodes. Eight evaluations were realized during the experimental period. Response surface analysis was used to evaluate the morphometric characteristics of the cladodes. Best water levels were between 55%.ETc and 65%.ETc and saline levels between 3.5 and 5 dS/m, maximizing the morphometric characteristics of Giant Sweet clone.
In this paper, a system of nonlinear equations for the maximum likelihood estimators as wel as the exact forms of the Fisher information matrix for Crovelli's bivariate gamma distribution and bivariate gamma beta distribution of the second kind are determined. An application of the results to the rainfall data from the city of Passo Fundo are provided.
The lack of error of experimental planning in agricultural field studies can result in rework, causing the waste of financial resources. The determination of the optimal size of the experimental plot for carrying out the treatments can minimize these problems. The objective of this paper was to estimate the optimal plot size for measuring reflectance in soybeans, without treatment, using the modified maximum curvature method and the maximum distance method. Reflectance readings were taken in the soybean crop with the aid of the GreenSeeker® equipment, in basic experimental units of 0.45 m², in an area of 7 lines and 8 meters in length. The data were collected in three phenological stages of soy (R4, R5.5 and R6), obtaining 63 simulations of experimental area in each stage. Based on the results, it is recommended to use plots of 7.20 m², with grouping of 4 lines of 4 m in length.
This article addressed traffic accidents in the urban perimeter of the city of Londrina- Paraná, Brazil, carried out in 2019. The information used was from the Integrated Emergency Trauma Care System and collected through the General Occurrence Registry (RGO) of the Fire Department of Londrina. For this study, accidents were selected in the collision category between vehicle drivers and motorized drivers. The study variables were the most frequent months and the accident codes defined by Integrated Trauma Emergency Service (SIATE). The article aimed to analyze the occurrences (collision) of vehicle drivers in 2019, seeking to make the population aware of accident prevention. The methodology adopted was descriptive, quantitative, obtained intentionally. The results show that these categories of drivers are responsible for more than 50% of the total accidents and more than 60% of the fatal victims. According to the results obtained, it is expected that the authorities may be sensitized, adopting measures that lead those responsible to more rigorous sanctions, thus making drivers aware of having greater responsibility when driving a vehicle, emphasizing that they need to redouble their attention and have responsibility in traffic, avoiding negligence and damage to the lives of others.
A Rasch Poisson counts (RPC) model is described to identify individual latent traits and facilities of the items of tests that model the error (or success) count in several tasks over time, instead of modeling the correct responses to items in a test as in the dichotomous item response theory (IRT) model. These types of tests can be more informative than traditional tests. To estimate the model parameters, we consider a Bayesian approach using the integrated nested Laplace approximation (INLA). We develop residual analysis to assess model t by introducing randomized quantile residuals for items. The data used to illustrate the method comes from 228 people who took a selective attention test. The test has 20 blocks (items), with a time limit of 15 seconds for each block. The results of the residual analysis of the RPC were promising and indicated that the studied attention data are not well tted by the RPC model.
Brazil is a major producer in the timber sector, mainly with the use of wood from species of the genus Eucalyptus, with 26.1% of planted forests located in Minas Gerais. Researchers and manufacturers have been searching for techniques with the objective of making full use of these forests, with a primary focus on greater growth. A modeling of growth curves is an alternative for the estimation of oral production andan important aid tool for the researcher's decision making. Growth curves are commonly studied by nonlinear regression models, which have important assumptions that if not met should be added to the model. The present work aims to select among nonlinear Logistic, Gompertz and von Bertalany regression models the most suitable to describe the growth in wood volume of Eucalyptus urophylla x Eucalyptus grandis hybrids in three Forest Site categories, including whether assumption deviations are required. Methods were executed by the Gauss-Newton iterative method implemented in nls() and it gnls() functions of the R software. Determination coecient, Akaike information criterion (AICc) and Residual Standard Deviation (RSD) were used as selection evaluators of the best model. The results demonstrate that for all site categories, the Gompertz model with addition of autoregressive parameters AR (1) is the most appropriate to describe the growth in wood volume of Eucalyptus urophylla x Eucalyptus grandis hybrids. The addition of the rst-order autoregressive parameter does not aect the quality of t, but it is the correct procedure. Site I, which presents the largest trees according to pre-dened variations, recorded 308 m3/ha of wood volume, followed by 286 m3/ha and 263 m3/ha for Sites II and III, respectively. The time for Site III to reach the maximum point of volume growth is between the fourth and sixth year, while the other sites are more precocious, reaching this point between the second and third year.
This article is a direct consequence of the authors’ desire to discuss the role of statistics in data analysis. The analysis of coronavirus (COVID-19) databases are used as to show simple, but powerful statistical frameworks. We do believe that models for assessing future trends in temporal data in general, and in cases and/or deaths of COVID-19, belongs to the area of (Bio)Statistics. Just as engineers use knowledge of physics, chemistry and often architecture, when constructing bridges, buildings and roads, statisticians use knowledge of mathematics, computer science and even physics for modelling, analysing, and forecasting in order to transform data into information. While the statistician’s contribution is rarely acknowledged, everyone knows that a building is a work of an engineer. Nonetheless, nowadays statistics has been gaining the attention that it deserves due to the rise of big data and data science that was built on the foundations of statistics. This article shows that, even with only basic knowledge of statistics, one can adequately collaborate with the community in dealing with very important issues such as the COVID-19 numbers. In order to model and to obtain predictions we use well-known distributions to statisticians working on survival analysis: gamma, Weibull and log-normal distributions. We also make use of singular spectrum analysis, a simple non-parametric time series methodology, for an analogous purpose. Survival analysis is a research area widely used in Biostatistics and even in Reliability, while time series analysis is widely used across areas where the data is measured along the time.
Coffee growing is one of the most important agricultural activities in the world market. Among the commercially relevant species, there is Coffea canephora,which can be divided into the varietal groups Conilon and Robusta. These varietal groups have complementary agronomic interests. Because of this, hybrids are obtained through the crosses between these groups. Given the difficulty in differentiating between two varietal groups genotypes in the field, the correct discrimination is essential for the definition of crosses in breeding programs. In this context, the objective was to apply a discriminant analysis (DA) to define functions to differentiate between varietal groups and hybrids of canephora, as well as to identify the most relevant phenotypic traits in these functions. Data from 165 genotypes from the Instituto Capixaba de Pesquisa, Assistência Técnica e Extensão Rural e do Centro Agronómico Tropical de Investigación y Enseñanza were used for which different plant traits were measured. The quadratic DA applied was the one with the best performance for genotype discrimination, with an average apparent error rate of 0.0333. Cercosporiose incidence, rust incidence and vegetative vigor were the most important traits in the varietal groups' discrimination.
Absenteeism is the practice or custom of an employee to be absent from workplace. Its causes are diverse and may aect the workers income as well as to cause operational disruption, stress the administration and also nancial losses for the company. Cluster analysis is a multivariate tool that can be used to determine groups in the sense that each group has its own characteristics in terms of the observed variables.In this sense, that technique can be used as a support to show which characteristics may contribute to absenteeism. We use the Ward hierarchical algorithm to build the clusters and to compare the groups the Kruskal-Wallis nonparametric test is adopted. Finally, a study on the strength of association among the variables is developed using Spearman's correlation and for the relationship among those variables related to absence and social aspects, we use the principal component analysis. Moreover, the study indicates the possibility to determine three heterogeneous groups in the company and to show characteristics in those groups which are potential factors that cause absenteeism to a greater or lower extent.
Statistical Process Control (SPC) stands out for the use of control charts and for repeatability and reproducibility (R&R) techniques. This work aimed at its applications in the aspects of pre-processing of structural monitoring. The experiment was carried out in a completely randomized design (CRD) with two sources of variation: eight aluminum beams with piezoelectric patches and five types of damage (D1 = baseline, D2 = 0.6g, D3 = 1.1g, D4 = 1.6g, D5 = 2.2g). All measurements were gathered at 30oC and with 20 repetitions for each condition case, producing a damage metric. In the R&R study, a low variation of repetition was observed (9.84%), but a high reproducibility (72.39%), representing that the damage metrics were similar for each situation, but a high variation among beams and damages. Based on this evaluation, the control charts helped to verify in which beams and damages these greatest variabilities were found. Concluding, the control charts for mean and individual measures as well as the R&R study were interesting tools for raw data pre-processing step for measurement error detection.
Animal behavior studies usually produce large amounts of data and a wide variety of data structures, including nonlinear relationships, interaction effects, nonconstant variance, correlated measures, overdispersion, and zero inflation, among others. We aimed to explore here the potential of generalized additive models for location, scale and shape (GAMLSS) in analyzing data from animal behavior studies. Data from 20 Romane ewes from two genetic lineages submitted to brushing by a familiar observer were analyzed. Behavioral responses through ear posture changes, a count random variable, and the proportion of time to perform the horizontal ear posture, a continuous random variable on the interval (0,1), with non-null probabilities in zero and one, were analyzed. The Poisson, negative binomial, and their zero-inflated and zero-adjusted extensions models were considered for the count data, whereas the beta distribution and its inflated versions were evaluated for the proportions. Random effects were also included to consider the multilevel structure of the experiment. The zero adjusted negative binomial model has better fitted the count data, whereas the inflated beta distribution performed the best for the proportions. Both models allowed us to properly assess the effects of social separation, brushing, and genetic lineages on sheep behavioral. We may conclude that GAMLSS is a flexible framework to analyze animal behavior data.
We develop best linear unbiased predictors (BLUP) of the latent values of labeled sample units selected from a finite population when there are two distinct sources of measurement error: endogenous, exogenous or both. Usual target parameters are the population mean, the latent values associated to a labeled unit or the latent value of the unit that will appear in a given position in the sample. We show how both types of measurement errors affect the within unit covariance matrices and indicate how the finite population BLUP may be obtained via standard software packages employed to fit mixed models in situations with either heteroskedastic or homoskedastic exogenous and endogenous measurement errors.
This work presents a cluster analysis approach aiming to determine distinct groups based on clinicopathological data from patients with breast cancer (BC). For this purpose, the clinical variables were considered: age at diagnosis, weight, height, lymph nodal invasion (LN), tumor-node-metastasis (TNM) staging and body mass index (BMI). Ward's hierarchical clustering algorithm was used to form specific groups. Based on this, BC patients were separated into four groups. The Kruskal-Wallis test was performed to assess the differences among the clusters. The intensity of the influence of variables on the prognosis of BC was also evaluated by calculating the Spearman's correlation. Positive correlations were obtained between weight and BMI, TNM and LN invasion in all analyzes. Negative correlations between BMI and height were obtained in some of the analyzes. Finally, a new correlation was obtained, based on this approach, between weight and TNM, demonstrating that the trophic-adipose status of BC patients can be directly related to disease staging.
Para obter a estimativa da variância aleatória em experimentos fatoriais completos e fracionados com dois níveis por fator avaliados sem repetições, Hamada e Balakrishnan (1998) fornecem uma lista de vários métodos. Assim, com base nessa revisão, o objetivo do presente trabalho consistiu em comparar as estimativas dos desvios-padrão com apenas influências das causas aleatórias de acordo com quatro métodos: de Lenth (1989), de Juan e Pena (1992), de Dong (1993) e sem nenhuma restrição aos dados, aqui denominado de desvio-padrão total. Para isso, foi simulada uma variável aleatória normal com 10.000 valores, cuja simulação foi repetida 16 vezes. Posteriormente, foram substituídos em cada um dos 16 conjuntos de dados, 0%, 1%, 2%, 3% e 4% dos valores aleatórios por outliers com o objetivo de quebrar a aleatoriedade da variável simulada. Com base na estimativa do erro percentual médio absoluto (EPMA) obtida em relação ao desvio-padrão aleatório paramétrico, concluiu-se, por meio da análise de regressão, que ela aumentou em função do aumento do percentual de substituição dos valores aleatórios por outliers, com exceção à obtida de acordo com o método de Juan e Pena (1992). Mesmo assim, para conjuntos de dados com até 3,68% de outliers, os melhores métodos de estimação do desvio-padrão aleatório (Saleatório) foram os de Lenth (1989) e de Dong (1993), por terem fornecido as menores estimativas do EPMA. Acima desse percentual e até 4% de outliers, o método de Juan e Pena (1992) mostrou-se ser melhor. No entanto, como a maior estimativa do EPMA proporcionada pelos três métodos de estimação foi muito baixa (4,00%), e ainda, como as diferenças observadas entre eles foram, praticamente, desprezíveis, concluiu-se que os três métodos forneceram boas estimativas do Saleatório e que, consequentemente, podem ser recomendados para estimar o quadrado médio do resíduo em experimentos fatoriais completos e fracionados com dois níveis por fator e com observações individuais por tratamento. Por outro lado, o método do desvio-padrão total não conseguiu evitar o efeito da não aleatoriedade sobre a estimativa do Saleatório.
In this paper, we proposed the Poisson-Weibull distribution for the modeling of survival data. The motivation to study this model since, in addition to generalizing the Weibull distribution, which is widely used in several areas of knowledge among them the Survival and Reliability analysis, it presents great exibility in the forms of the hazard function. The Poisson-Weibull distribution was created in a composition of discrete and continuous distributions where there is no information about which factor was responsible for the component failure, only the minimum lifetime value among all risks is observed. The maximum likelihood approach was used to estimate the parameters of the model. Also was conducted a simulation study to examine the mean, the bias, and the root of the mean square error of the maximum likelihood estimates of the proposed model according to the censoring percentages and sample sizes. The model selection criteria were also applied, in addition to graphic techniques such as TTT-Plot and Kaplan-Meier. Application to the real data set was used to illustrate the usefulnessof the distribution.
This study aimed to determine the size and shape of experimental plots that provide maximum precision using relative information method. This trial was conducted at the Federal Institute of Bahia. Plant height, cladode length, cladode width, cladode thickness, cladode area, cladode area index, number of cladodes, cladode total area and yield were measured in the third production cycle, 930 days after planting. The plants, defined as basic units, were arranged in 39 plot sizes so that the crop would fill the whole experimental area. Then, plot shapes with higher relative information and equal plot size in basic units were selected. The experimental plot with eight basic units in size ensures higher efficiency in the experimental evaluation. This combination between size and shape, besides meeting all evaluation requirements of the characteristics normally assessed in studies with forage cactus pear, has the maximum control of soil heterogeneity, thereby decreasing experimental error and significantly increasing precision.
Apesar dos lucros auferidos no segmento de saúde suplementar, o número de segurados vinculados a ele oscilou entre 2015 e 2018, em contraste com o comportamento apresentado entre 2000 e 2014, quando só cresceu. Para compreender parte desse fluxo, isto é, a saída dos segurados desse mercado, objetiva-se analisar o tempo de permanência do segurado em planos de saúde, a partir de dados compostos por 122.381 segurados (e ex-segurados) acompanhados entre os anos de 1984 e 2018. Utilizando-se da análise de sobrevivência tradicional, por meio do estimador de Kaplan-Meier e de modelos paramétricos e semiparamétricos, destacam-se os seguintes resultados: a) a mediana do tempo de permanência no plano é de 4,62 anos; b) a massa de seguradoras é composta (ao longo dos anos observados) predominantemente por mulheres, solteiras, jovens, titulares e aderentes ao contrato de individual/familiar; c) conforme o modelo de Cox selecionado, ser homem (em relação à mulher), ser jovem (em relação ao adulto), ser dependente (em relação ao titular) e ser casado (em relação ao amasiado) aumentam o risco de saída da operadora analisada. Espera-se que esses resultados auxiliem a operadora analisada a (re)direcionar suas políticas comerciais e de subscrição de riscos.
Social inequality is the phenomenon that differentiates between people in the context of the same society, placing some individuals in structurally more advantageous conditions than others. It manifests itself in all aspects: political, economic among others. The main causes of inequality are investment lack in social areas, health and education. Among the consequences of inequality, we highlight: increased violence, poverty, delay in economic progress; hunger, destruction and infant mortality; young marginalization people, and finally; rising unemployment. Among the main inequality types, we highlight: people with and without disabilities, regions, races; income and sex. To measure this inequality, we highlight HDI, Theil and MPI. A person with a disability is any person who presents a loss or abnormality that generates an inability to perform one or more activities, and these characteristics hinder their social inclusion, access to the labor market, transportation, education, financing and training; urban and environmental barriers, and finally; ignorance of employers. Situations like these provide disabilities people with lower wages when employed, worse purchasing power, less social participation providing greater exclusion and disadvantaged situations when compared to those without disabilities. For this work we used exploratory analysis techniques considering data sets from the 2010 IBGE Census and UNDP.
The COVID-19 pandemic has spread rapidly around the world in a frightening way. In Brazil, the third country with the highest number of infected and deaths from the disease, it is important for government health authorities to identify the federation units that stand out in cases and deaths from this disease to target resources. The circular scan statistic proposed by Martin Kulldorff allows to identify with some statistical significance the units of the federation that stand out in relation to the number of cases and deaths of COVID-19 in Brazil. Such units of federation are known as clusters. Once these clusters were identified, we used the coefficients of incidence and lethality to better describe the behavior of these clusters during three phases of the pandemic: the initial phase, the peak phase, and also the stability and fall phase. We observed changes in the location of the clusters identified in these three phases and used the R software and also the SaTScan software to obtain the maps and results, which were consistent with what was reported by the Brazilian media.
Breast cancer is one of the most common diseases among women worldwide with about 25% of new cases each year. In Brazil, 59,700 new cases of breast cancer were expected in 2019, according to the Brazilian National Cancer Institute (INCA). Survival analysis has been an useful tool for the identifying the risk and prognostic factors for cancer patients. This work aims to characterize the prognostic value of demographic, clinical and pathological variables in relation to the survival time of 2,092 patients diagnosed with breast cancer in Parana State, Brazil, from 2004 to 2016. In this sense, we propose a Bayesian analysis of survival data with long-term survivors by using Weibull regression models through integrated nested Laplace approximations (INLA). The results point to a proportion of long-term survivors around 57:6% in the population under study. In regard to potential risk factors, we namely concluded that 40-50 year age group has superior survival than younger and older age groups, white women have higher breast cancer risk than other races, and marital status decreases that risk. Caution on the general use of these results is nevertheless advised, since we have analyzed population-based breast cancer data without proper monitoring by a healthprofessional.