Tobacco is a globally cultivated crop featuring distinct quality variations among leaves from different geographical origins. To develop a rapid, robust, and accurate method for multi-origin traceability, this study employed near-infrared spectroscopy combined with rapid chemical composition analysis to obtain 70 chemical components in samples from nine major tobacco-producing regions in China and four other countries (the United States, Brazil, Zimbabwe, and Zambia). One-way analysis of variance (ANOVA) and hierarchical cluster analysis (HCA) were used to investigate regional chemical differences. Discrimination models were built using a support vector machine (SVM), a backpropagation neural network, and a random forest. The best model was interpreted using permutation feature importance (PFI) to identify key markers for origin discrimination. One-way ANOVA revealed significant differences (p ≤ 0.001), and HCA demonstrated clear regional patterns. The SVM-hybrid kernel achieved the best performance with 97.96% test accuracy and macro-average recall, precision, and F1 scores of 0.9836, 0.9806, and 0.9821, respectively. The PFI algorithm was employed to identify and rank the top 20 key chemical components influencing the geographical origin discrimination. The top ten key components were Fru-Asn, succinic acid, rutin, Fru-Val, sulfate, serine, phosphate, starch, potassium, and Fru-Gly. This study integrated chemometrics, near-infrared, rapid chemical analysis, and interpretable machine learning to accurately distinguish tobacco origins, reveal regional traits, and offer insights into geographical traceability and chemical profiling.
To address complex chemical relationships and geographical sample imbalance in tobacco origin classification, a method combining chemical composition imaging with a two-dimensional convolutional neural network (2D-CNN) and the synthetic minority over-sampling technique (SMOTE) with threshold moving is proposed. In this approach, multidimensional chemical components are transformed into structured 2D images. Then, a 2D-CNN deep learning model is constructed to capture intricate correlations among chemical indicators through two-dimensional convolutions. Classifier bias arising from class imbalance is mitigated by combining SMOTE with threshold-moving techniques. The results show that the 2D-CNN classification model achieved an overall accuracy of 0.9764 on the test set, with an average precision of 0.9477, a recall of 0.9511, and an F1-score of 0.9492 across eight ecological areas, indicating high model performance. Under the same imbalance handling, the 2D-CNN outperformed a one-dimensional CNN (1D-CNN) by 2.51% in average F1-score, confirming that chemical composition imaging effectively extracts complex inter-indicator relationships. The integration of SMOTE and threshold moving effectively alleviates the impact of class imbalance, significantly enhancing the recognition rates for minority areas. Furthermore, to independently validate the effectiveness of the imbalance handling strategy, it was applied to a separate public image dataset (the Niphad Grape Leaf Disease Dataset). Compared to the baseline model without any imbalance mitigation, the absolute recall of the minority class increased by 40 percentage points.
IntroductionTo identify the key chemical components affecting the sensory irritation of tobacco leaves, a binary classification model for high and low irritation was constructed based on 78 chemical components and sensory evaluation scores of 353 tobacco leaf samples.MethodsFirst, the median absolute deviation (MAD) method was applied to remove outliers from the high- and low-irritation samples. Then, the ReliefF algorithm was applied for dimensionality reduction, selecting 33 core features to eliminate data redundancy. Using the selected features, a random forest (RF) algorithm was employed to build the classification model, and the optimal number of decision trees was determined to be 60.ResultsThe ReliefF-RF model achieved an accuracy of 84.38% on an independent test set, with precision, recall, and F1-score all at 86.49%, outperforming the original RF model as well as other machine learning models such as support vector machine (SVM) and k-nearest neighbors (KNN). Through feature importance evaluation, eight key chemical indicators were identified: total nitrogen, total alkaloids, cryptochlorogenic acid, oleic acid + linolenic acid, reducing sugar, sugar-nitrogen ratio, Fru-Asp, and neochlorogenic acid.DiscussionSHapley Additive exPlanations (SHAP) analysis revealed that higher levels of nitrogenous compounds were strongly associated with increased irritation, whereas elevated levels of sugar components, specific organic acids, and amino acid derivatives were associated with reduced irritation. Notably, Fru-Asp exhibited a complex non-linear response, where both extremely high and low levels contributed to higher irritation. This study provides a useful reference and data support for the targeted regulation of cigarette irritation.
This dataset comprises 347 complete spectra of tobacco leaves, acquired using an Antaris II Fourier Transform Near-Infrared spectrometer equipped with an integrating sphere diffuse reflectance sampling system. The samples were collected from six countries: Argentina, Brazil, Zimbabwe, the United States, Tanzania, and Zambia. All samples were dried using an FD240 oven (Binder GmbH, Germany), ground using a ZM200 grinder (Retsch GmbH, Germany), and sieved through a 0.250 mm mesh screen. NIR spectra were acquired with the following parameters: a spectral range of 4000–10000 cm−1, a resolution of 8 cm−1, and 64 scans. The dataset was partitioned into a training set (70%) and a validation set (30%) using stratified sampling. Six preprocessing techniques were applied: Savitzky-Golay (SG) smoothing, Multiplicative Scatter Correction (MSC), Standard Normal Variate (SNV), First Derivative (1D), Second Derivative (2D), and Mean Centering. Partial Least Squares (PLS) regression was utilized to establish predictive models for 13 chemical indicators: total alkaloids, reducing sugars, total sugars, total nitrogen, K, Cl, pH, starch, neochlorogenic acid, chlorogenic acid, cryptochlorogenic acid, scopoletin, and rutin. The Random Forest (RF) algorithm was employed to create the origin classification model. Ultimately, through the selection of appropriate spectral preprocessing methods, the quantitative prediction models established using this dataset all achieved coefficients of determination (R2) exceeding 0.83, demonstrating robust predictive performance. Furthermore, the geographical origin classification models yielded an overall validation accuracy greater than 0.8, indicating strong classification performance. Consequently, this dataset is confirmed to be accurate and reliable, capable of providing foundational data for other researchers to construct near-infrared (NIR) spectral databases and develop NIR prediction models. The developed models exhibit excellent predictive capabilities and can be utilized in practical applications as an alternative to classical chemical analysis methods for these chemical indicators.
Tobacco leaf position is closely associated with its quality whose material basis is the chemical components of tobacco leaf. In recent years, near-infrared (NIR) spectroscopy combined with algorithmic models has emerged as a popular method for identifying the tobacco leaf position. However, when applied to leaf position discrimination, these models often rely on principal components derived from dimensionality-reduced spectral signals, resulting in limited interpretability and difficulty in identifying key chemical components. Chemical composition data combined with algorithmic models can also be used to discriminate tobacco leaf positions. However, the acquisition of chemical components relies on traditional instrumental analytical methods. As a result, the acquisition of chemical composition data is time-consuming and labor-intensive, involving only a limited number of compounds. The study proposes a novel approach that integrates machine learning with advanced interpretability techniques for both tobacco leaf position discrimination and analysis. Based on the 70 tobacco leaf chemical components obtained using near-infrared rapid analysis technology, tobacco leaf position discrimination models were built using Support Vector Machine (SVM), Back Propagation Neural Network (BPNN), and Random Forest (RF). Particle swarm optimization (PSO) was used to optimize parameters of each model. Chemical components were analyzed for statistical significance across leaf positions, and their influence on model predictions was interpreted using SHapley Additive exPlanations (SHAP). The experimental results showed that among all models, the SVM- hybrid kernel demonstrated the most robust and accurate performance, achieving discrimination accuracies of 98.17% and 96.33% on the training and test sets, respectively. SHAP analysis provided a clear ranking of feature importance and revealed the positive and negative contributions of individual chemical components. The proposed method can be useful for position traceability and chemical feature analysis of various crops.
The trace detection of both aqueous and gaseous nicotine is significant for health monitoring and environmental analysis. In this study, a novel covalent organic framework (COF)-based composite was designed to enrich and detect nicotine using surface-enhanced Raman scattering. The negatively charged TpPa-SO3H exhibited exceptional enrichment capabilities for nicotine, demonstrating an adsorption capacity of 148.0 mg center dot g-1. Chemical interactions, including it-it stacking and acid-base interactions, were observed between the COFs and nicotine molecules. Furthermore, chemical enhancement induced by electron transfer between TpPa-SO3H and nicotine was corroborated through energy level analysis. By leveraging the bimetallic synergism between core-shell Au and Ag to achieve electromagnetic enhancement, the TpPa-SO3H-Au@Ag composite achieved a detection limit of 6.0 x 10-10 mol center dot L-1 for aqueous nicotine. The detection of nicotine in e-cigarette oil and Cambridge filter demonstrates its application capability for different kinds of real samples. By the crosslinking with chitosan, the synthesized aerogel was successfully applied for the on-site gaseous nicotine detection in cigarette smoke. Additionally, the plasmonic aerogel enabled the identification of different kinds of cigarettes using artificial intelligence-assisted gaseous fingerprint spectra recognition and classification.
Full-wavelength near-infrared (NIR) spectroscopy faces significant challenges due to the strong collinearity among spectral variables and the presence of variables that are highly sensitive to sample fluctuations. Additionally, not all spectral variables contribute equally to the NIR model. Weakly influential variables, although not important on their own, can provide substantial improvement when combined with stronger variables, thus increasing both model stability and prediction accuracy. Therefore, this study proposes a new variable selection method called outlier removal with weight penalization and aggregation (OR-WPA). The method begins by removing outlier spectral variables with high coefficient of variation, which enhances model stability. During the variable selection process, multiple submodels are constructed based on variable subsets, with variable weights assigned according to the absolute values of regression coefficients. A moving window is applied to average the weights, and variables with excessively high weights are penalized, promoting the selection of weakly influential variables that positively contribute to model accuracy. The variable space is iteratively reduced, and the subset of variables associated with the highest predictive accuracy is selected as the final characteristic variable combination. The OR-WPA method was evaluated on three NIR spectral data sets, involving corn, heated tobacco substrate, and flue-cured tobacco. The results were compared with three advanced variable selection methods: Monte Carlo uninformative variable elimination, competitive adaptive reweighted sampling, and bootstrapping soft shrinkage. The results indicate that OR-WPA demonstrates better predictive performance, particularly in predicting low-content components, where it significantly enhances both the accuracy and stability of the NIR model.
The strong absorption characteristics of water in the near-infrared (NIR) spectral region significantly hinder the quantitative analysis of chemical compounds. Using the quantitative analysis of Amadori compounds in tobacco leaves as a case study, this study proposed a spectral decomposition optimization algorithm (SDOA) to mitigate water interference in NIR spectra, which helped improved the accuracy of predictive models. The method utilized a reconstructed spectral matrix derived from aged tobacco leaves across 13 moisture levels, ranging from 4 % to 16 % water content. By constructing a difference spectrum matrix, performing singular value decomposition, and applying projection transformations to the original NIR spectra, the influence of water on the spectral was reduced. A partial least squares (PLS) model was established based on the moisture-corrected spectra to predict the content of 17 Amadori compounds in tobacco. The results demonstrated that: (1) After SDOA processing, the average correlation coefficient (R-_2) of spectra at different moisture levels increased from 0.9903 to >0.9999, indicating significantly enhanced spectral consistency; (2) Among the 17 Amadori compounds, Fru-Pro, which was present at a relatively high concentration in tobacco leaves, exhibited minimal water interference. For the remaining 16 compounds, the SDOA-PLS models achieved robust predictive performance, as evidenced by evaluation metrics including the coefficient of determination (R-2), root mean square error (which ranged from 0.8705 to 0.9859), and relative prediction deviation (which ranged from 2.778 to 8.423). (3) One-way analysis of variance confirmed no significant differences (alpha = 0.05) in the predicted Amadori compound contents across tobacco samples with varying moisture levels using the SDOA-PLS method. The SDOA-PLS model effectively eliminates water interference in NIR spectra and significantly improves the accuracy and stability of Amadori compound content prediction in tobacco leaves.
A colorimetric test strip was developed for the qualitative and quantitative analysis of Boric acids in water-based adhesive. The influence of various experimental conditions was examined to optimize the analytical performance of test strip. The test strip for the detection of boric acid was carried out using 0.5mmol 5-hydroxy flavones and 2.6mmol citric acid, temperature at 80℃ and 40% nitric acid. The developed test strip is rapid within 6 min. A yellow complexation is observed in the test strip when boron acid and 5-hydroxyflavone is present due to the complex reaction in citric acid and nitric acid. The reaction was proportional to its concentration of boric acid ranging from 25 to 1000 mg/L, with the detection limit of 8 mg/L. With its simple and effective procedures, the developed test strip showed a great promising for on-site detection of boric acid in water-based adhesive.
Aerosol pollutants significantly cause health concerns. Herein, we established an original real-time aerosol exposure system that used a self-designed bionic-lung microfluidic chip. The chip features a 4 x 4 intersecting array within gas and liquid layers, creating 16 distinct microenvironments. A membrane situated between the layers offers attachment for cells and establishes a gas-liquid interface. This design provides a reliable screening capacity for investigating the biological effects of aerosol exposure in vitro by manipulating the gas and/or liquid conditions. Using this system, we validated that cigarette smoke (CS) aerosol triggered a concentration- and timedependent reduction in cell viability and intracellular glutathione levels, accompanied by an increase in intracellular reactive oxygen species and Fe2+. Furthermore, CS aerosol significantly downregulated the expression of GPX4, SLC7A11, and FTL mRNA while inducing a notable increase in that of ACSL4 mRNA. Additionally, CS aerosol markedly stimulated the release of proinflammatory cytokines. Crucially, the ferroptosis inhibitor deferoxamine mesylate reversed these biological indicators. These results demonstrate that our novel bionic-lung chip presents a suitably achievable approach to investigate the biological effects induced by aerosol exposure.
In this study, we investigated the additive patterns observed in near-infrared (NIR) diffuse reflectance spectra within the context of food and pharmaceutical product formulation. Employing the Kubelka–Munk theory, we examined the linear correlation between spectra and concentration in polymer materials and tobacco powder samples. Our findings confirm the principle of spectral additivity in diffuse reflectance spectra, demonstrating a linear relationship when the sample scattering coefficient remains constant. Moreover, our results validate the feasibility of substituting actual mixed spectra with NIR additive spectra in tobacco leaf systems. This approach can potentially enhance formulation design, thereby improving efficiency and accuracy while expanding the scope and combination of formulation materials. Furthermore, this study offers a rapid, information-rich, and environmentally friendly alternative to traditional methods, with significant implications for the future of product formulation.
Near-infrared (NIR) spectroscopy has gained wide acceptance across various fields as a result of advances in portable equipment that can record spectra on site or at production lines. Continuous wavelet transform (CWT) can transform traditional one-dimensional (1D) NIR spectra into more informative two-dimensional (2D) spectrograms, thus enhancing the analysis and interpretation of spectral information. This study introduces a high-efficiency 2D CWT-EfficientNetV2 regression model to optimize NIR spectroscopy applications. A novel progressive screening strategy is employed to select the optimal wavelet functions and scales for CWT, which are then used to transform the features into wavelet coefficient matrices. Direct digital mapping (DDM) with Gray colormap generates 2D spectrograms from matrices, significantly preserving the representation of wavelet coefficients. The 2D CWT-EfficientNetV2 model was used to predict the content of five polyphenols in tobacco leaf samples with superior performance compared to partial least squares regression (PLSR) and other high-efficiency models. Moreover, to further validate the robustness and reliability of the proposed method, two additional public NIR spectral datasets were included in this study. The model achieves lower root mean square error of prediction (RMSEP), as well as higher coefficient of determination of prediction (RP2) and the ratio of the standard error of prediction to the standard deviation of the reference values (RPD) on the test datasets. These results demonstrate that the 2D CWT-EfficientNetV2 model is a robust and efficient approach for the accurate quantification of various target compounds utilizing NIR spectroscopy.
Tobacco is one of the most widely cultivated non-food cash crops worldwide, and the quality of tobacco procured from different geographical locations varies considerably. This study proposes a comprehensive strategy for investigating tobacco quality using near-infrared (NIR) prediction models with moisture-adaptive corrections. This strategy enables the rapid and efficient quantification of 70 chemicals in tobacco by reducing the effect of moisture on the NIR spectra of tobacco samples. Additionally, this strategy has been proposed for the geographical discrimination and part identification of tobacco samples, with the Mahalanobis distance analysis, the accuracy of predicted values is higher than 81.5%. This study confirms that the tobacco chemical composition from different regions in China is inconsistent.
The content of nicotine, a critical component of tobacco, significantly influences the quality of tobacco leaves. Near-infrared (NIR) spectroscopy is a widely used technique for rapid, non-destructive, and environmentally friendly analysis of nicotine levels in tobacco. In this paper, we propose a novel regression model, Lightweight one-dimensional convolutional neural network (1D-CNN), for predicting nicotine content in tobacco leaves using one-dimensional (1D) NIR spectral data and a deep learning approach with convolutional neural network (CNN). This study employed Savitzky–Golay (SG) smoothing to preprocess NIR spectra and randomly generate representative training and test datasets. Batch normalization was used in network regularization to reduce overfitting and improve the generalization performance of the Lightweight 1D-CNN model under a limited training dataset. The network structure of this CNN model consists of four convolutional layers to extract high-level features from the input data. The output of these layers is then fed into a fully connected layer, which uses a linear activation function to output the predicted numerical value of nicotine. After the comparison of the performance of multiple regression models, including support vector regression (SVR), partial least squares regression (PLSR), 1D-CNN, and Lightweight 1D-CNN, under the preprocessing method of SG smoothing, we found that the Lightweight 1D-CNN regression model with batch normalization achieved root mean square error (RMSE) of 0.14, coefficient of determination (R2) of 0.95, and residual prediction deviation (RPD) of 5.09. These results demonstrate that the Lightweight 1D-CNN model is objective and robust and outperforms existing methods in terms of accuracy, which has the potential to significantly improve quality control processes in the tobacco industry by accurately and rapidly analyzing the nicotine content.
为研究卷烟主流烟气中与感官相关的酰胺成分,基于分散固相萃取技术结合GC-MS/MS同时测定卷烟主流烟气中21种酰胺类化合物.通过优化提取、净化、色谱分离和质谱多反应监测等关键前处理条件和仪器参数,建立了酰胺类化合物的分析方法;对12个不同品牌和不同焦油释放量的市售卷烟进行测试,研究了卷烟烟气中21种酰胺类化合物的释放量水平及其影响因素.结果表明:①选择石墨化炭黑作为分散固相萃取吸附剂,烟气基质净化效果好、21种酰胺类化合物的回收率高,所建分析方法灵敏度高、准确性好、抗基质干扰能力强.②酰胺类化合物沸点较高,在卷烟主流烟气中主要分布于粒相物中.③卷烟样品中可检出其中的16种酰胺类化合物,释放量范围为0.4~303.9 ng/mg(以平均单位TPM释放量计);其余5种未检出.④12种酰胺类化合物的释放量与TPM高度正相关(相关系数大于0.8).⑤T检验结果显示,7种酰胺类化合物在混合型卷烟中单位TPM释放量与烤烟型卷烟有极显著性或显著性差异,混合型卷烟明显高于烤烟型卷烟.
To study the effect of heating temperature on the dynamic distribution of aerosol from the tobacco section of heated tobacco products, heated tobacco product samples were smoked with the same matched heating apparatus at different heating temperatures. An in-situ aerosol characterization system was used for sampling and on-line detection of aerosol from each detection point in the tobacco section.The dynamic distributions of the main aerosol physical properties formed at 300, 330, and 360 ℃heating temperatures were investigated during smoking. Furthermore, the temperature fields inside the tobacco section heated at different heating temperatures were characterized by thermocouples to clarify the influence of temperature distribution on aerosol distribution. The results showed that: 1)The particle size of aerosol released from the tobacco section to the cavity section at the three heating temperatures showed approximately normal logarithmic distribution. The particle size and particle number concentration of the aerosol increased with increasing heating temperature. 2)The distributions of count median diameter, particle number concentration, and volume concentration were basically the same at the different heating temperatures. The aerosol with larger count median diameter and higher number concentration and volume concentration was mainly distributed in the middle and front sections near the heating element. 3)The distribution area of larger count median diameter and higher particle number concentration and volume concentration in the tobacco section increased significantly with the increase of heating temperature. 4)The temperature distribution in the tobacco section of the heated tobacco product directly correlated with the aerosol distribution. The higher temperature was conducive to aerosol formation due to the presence of high-concentration vapor components.
Removal of 1,3-butadiene from cigarette smoke plays an important role in human health and environmental protection. Herein, a series of UiO-66 X% containing different ratios of the -NH2 group was synthesized via the solvothermal method by using terephthalic acid (H2BDC) and 2-aminoterephthalic acid (NH2-BDC) as ligands. Using GO as support, a series of UiO-66-NH2/GO Y% were prepared by controlling the ratio of UiO-66-NH2 and GO. The effects of -NH2 and GO contents on the structure and composition of MOFs were investigated. Finally, the different -NH2 contents of UiO-66 X% and the different GO contents of UiO-66-NH2/GO Y% were applied in 1,3-butadiene removal from cigarette smoke. The results showed that UiO-66 X% with the higher contents of -NH2 showed a higher rate of 1,3-butadiene removal, and UiO-66-NH2/GO Y% with the GO contents of 5% showed the highest removal rate of about 33.85%, which was 25.54% higher than that of activated carbon. In addition, the saturation capacity of the adsorbent materials for 1,3-butadiene was as high as 210.01–239.54 mg/g, showing great potential in reducing harmful components in cigarette smoke and environmental protection.
Potential human health and environmental issues associated with aerosol release as a result of biomass pyrolysis have received a widespread attention. Accurate aerosol sampling at source can help to unravel the complex mechanisms involved. This paper describes a novel aerosol microprobe sampling system combined with a smoking cycle simulator and a differential mobility spectrometer (microprobe-SCS-DMS) for spatially resolved in-situ characterization of aerosol release inside a heated tobacco product. A purposely made microprobe was accurately positioned to sample the aerosol generated at different positions within the heated distillation and pyrolysis zone. The aerosol particle behaviour at source was studied under a combined heating and air-flow cycle. Detailed measurements on the aerosol formed from both the axial and radial directions inside the heated tobacco rod were obtained, including the aerosol particle size distribution, particle number concentration, and particle volume concentration during external air flow perturbation. The results provided detailed information on the formation, diffusion, transfer and filtration of the aerosol which would otherwise not be possible. The microprobe-SCS-DMS system described in this work could be useful for other similar challenges involving aerosol in-situ characterization.
Most studies have focused on the pulmonary toxicity of inhaled PAHs to date; therefore, their hepatotoxic consequences are yet unknown. The main aim of this study is to examine the association between urinary polycyclic aromatic hydrocarbons (PAHs) and liver function parameters among the US population. The data included in this study were from the National Health and Nutritional Examination Survey (NHANES) 2003–2016. Finally, we included 2515 participants from seven cycles of the NHANES. Logistic regression was performed to calculate the association between each PAH and liver function parameters (elevated vs. normal) with odds ratio (OR) and 95
为研究抽吸参数对电加热卷烟气溶胶粒数和粒径的影响,采用吸烟循环模拟机-快速粒径谱仪测试电加热卷烟气溶胶的粒数浓度和粒数分布,考察了抽吸容量(35、55、75 mL)、抽吸持续时间(2、3、4 s)和抽吸间隔(30、60 s)对电加热卷烟气溶胶粒数浓度和粒数中值粒径的影响.结果表明:①在单口方波抽吸条件下,电加热卷烟气溶胶的粒数浓度随抽吸过程实时降低,粒数中值粒径随抽吸过程实时变小;②抽吸容量和抽吸间隔的增大可明显增加气溶胶的粒数浓度;当抽吸容量增加至75 mL时,粒数中值粒径降低;抽吸间隔对粒数中值粒径无影响.③抽吸持续时间对电加热卷烟气溶胶粒数浓度和粒数中值粒径的影响均较小.