To look for factors of the COVID-19 spreading in the whole world currently, an empirical study has been tried by using a multi-regression analysis for mortality rates of 47 prefectures as an objective variable, and various indices as the explanatory variables. A support vector machine method was applied to deal with a nonlinear relationship between objective and explanatory variables, and a sensitivity analysis was applied to search the factors of the COVID-19 mortality. Welfare, urbanization, poverty rate, service industry, and sex ratio were obtained as dangerous factors which increase mortality, while single-person households, meals, and sleep were obtained as defensing factors which decrease mortality. Novel and useful knowledge for prevention measure of the COVID-19 was obtained: three factors of urbanization, service industry, and single-person household relating to the Three Cs contribute largest to the mortality, and two factors of welfare and poverty rate, reflecting the reality' of the poor people also contribute.
近年のわが国の重大社会問題の一つである自殺には地域差が存在するため,都道府県別の自殺死亡率に有意な影響を与える要因の解明を目的とする実証研究を試みた.47都道府県の男女別年齢調整自殺死亡率を目的変数,それとの関連が推測される健康,経済,社会,自然分野の指標54種を説明変数としてサポートベクター回帰分析を行い,自殺死亡率に対する決定要因を探索し,その相対的影響度を推定した.その結果,男女別にそれぞれ12種の要因が得られ,男性では精神保健福祉士数,家計収入,患者数などの要因,女性では悩み相談,出生率,残業時間などの要因の影響が大きいことを見出した。また,これまで未検証の精神保健福祉士数や残業時間,精神状態が有意の影響を与えるが,自殺率との関係が深いとされてきた失業率や離婚率は決定要因にはならなかった.さらに,自殺率が最も高い秋田県について決定要因の結果に基づき自殺対策の提言を試みた.
人間活動にともなう地球の持続可能性を評価する指標の1つに,エコロジカル・フットプリント(EF)がある。本研究では,47都道府県のエネルギーと資源の消費データ基づき,グローバル・フットプリント・ネットワーク(GFN)が算出した2010年の日本のEF値を各都道府県に割り当てることを試みた。都道府県のEFは,世界自然保護基金(WWF)の定義によるカーボン・フットプリント,耕作地,牧草地,漁場,森林地,生産能力阻害地の6つのカテゴリーに分けて計算した。その結果,鉄鋼や石油化学などの重化学工業の立地県(山口,大分,岡山,広島,和歌山など)がEF値で上位5県を占めるが,これらの県は都道府県別の総環境負荷量(EF×都道府県人口)でみると,人口規模が大きくないため9位以下になり,東京,神奈川,愛知,大阪などの大都市圏が上位を占めた。さらに,他地域での都市活動等が影響する産業によるエネルギー消費(CO2排出量)を47 都道府県の平均値を用いてEFを再評価した。そのEF値でみると,都道府県間の相違は減少したが,民生からのCO2排出量が多い北海道・東北地方,東京,中国地方,沖縄等のEF値が全国平均より高い傾向が認められた。その各都道府県のEF 値に都道府県人口を乗じて求めた総環境負荷量では,東京,神奈川,大阪,愛知,埼玉,北海道,千葉を初めとする人口規模の大きな都道府県で総環境負荷量が多いことを明らかにできた。
近年,国家間あるいは各国内の所得格差が幅広い関心を集めている.格差の原因に関する理論的研究はこれまでに数多く行われているが,現実の所得分布を再現する理論モデルは未だ得られていない.所得格差の原因を解明する目的で,所得分配に影響すると考えられる要因を説明変数として回帰分析(OLS)を行い,決定要因を探索した実証的研究が多数報告されている.しかし,これまでは少数の説明変数と線形のOLSのために高い精度のモデルが得られていなかった.本研究では,非線形回帰分析手法の一つであるサポートベクターマシン(SVM)を用いて161カ国のジニ係数(目的変数)と経済,政治,教育,健康,技術分野の57種の説明変数との相関を解析し,感度分析法を用いて57種の変数の中から所得格差の決定要因を探索する実証研究を試みた.その結果,25種の要因によって161カ国のジニ係数を決定係数(R2)0.795という高い精度で再現するモデルを構築した.また,25種の決定要因の中では政治的要因の寄与が最大であり,次いでGDP等の経済的要因と医療費等の健康要因がほぼ同程度の寄与であることが判明した.
A large-scale empirical experiment to analyze determinants of national happiness of human well-beings has been carried out based on happiness data of nations in the world as dependent variable and numerous factors as explanatory variables. The correlation between the happiness data containing World Database of Happiness for 149 countries and the data of 56 factors of countries’ indices such as economical, political, social, health, resource, environmental, life-style, and cultural fields was statistically analyzed to evaluate the influence of each factor on happiness. The determinants of happiness across nations were investigated by training non-linear regression support vector machine (SVM) models using the data of happiness and the country factors, and by optimizing the explanatory variables by the sensitivity analysis method. The results indicate that 20 factors satisfactorily represent the happiness data of 130 countries with the root mean squared error of 0.48 and the coefficient of determination of 0.867. A nonlinear regression technique like SVM is crucial for constructing a happiness predicting model due to the high nonlinear relationship between happiness and the explanatory variables. It was also revealed that health is the most important among various factors which influence the happiness due to the large contribution of health factors (e.g. life expectancy and mortality rate) to happiness. It is suggested that the direct contribution of economical factors such as gross domestic product (GDP) to happiness is not significant, but their indirect effect is not negligible through in the health condition.
格付け会社が行っている国債の格付けを統一的に再現するモデルを作成するために,2011年4 月末時点での158カ国の国債格付けと44種の経済・政治指標との相関をサポートベクター回帰 (SVR)により一括解析する大規模実験を行った.感度分析により有効な指標を選択し,SVRを最適化した結果,格付けの予測値と設定値のよい一致(平均二乗誤差0.070,決定係数0.925)が得られ,格付け会社が行っている国債格付けを公開の数値データのみから高精度で再現するモデルを作成することができた.解析に用いた各種指標の中では政治指標の寄与がきわめて高いことから,格付け会社が政治的要因を重視していることが判明した.我が国の国債の格付けについては,国内の格付け会社が高目の評価を行っているのに対し,海外の格付け会社は我が国の債務超過を大きく評価して厳しい格付けをしている傾向が明らかになった.格付けと指標との相関は非線形性が高いことから,国債の格付けを再現するモデルを作成するためには,感度分析により有効な指標を選択し,SVR等の非線形解析手法を用いることが不可欠であることが分かった.
Recently, ecological footprint (EF) receives much global attention as a measure to estimate the human impact on the earth. However, the evaluation of EF values is not so easy because of the complexity of its process and the need of numerous data. For that purpose it is required to develop a convenient model for estimating EF values for a variety of countries based on easily available data such as gross domestic product (GDP) and others. A large-scale regression experiment to analyze the comprehensive determinants of EFs across many countries has been carried out. A nonlinear support vector machine (SVM) method was applied to the regression analysis between EFs (dependent variable) of 162 countries and 32 factors (explanatory variables) in various fields. Optimum factors for the modeling were determined by using the sensitivity analysis method as a variable selection technique in SVM. It is demonstrated that 19 factors satisfactorily reproduce the EFs of 162 countries with a coefficient of determination (R2) of 0.930, which is remarkably superior to those of the ordinary least squares (OLS) method. It also is revealed that various factors such as meat consumption and air pollution as well as geographical factors such as population and land area, and economic factors such as GDP and GDP per capita have to be taken into account to construct a model for estimating EFs with high accuracy.
The Carcinogenicity Reliability Database (CRDB) was constructed by collecting experimental carcinogenicity data on about 1,500 chemicals from six sources, including IARC, and NTP databases, and then by ranking their reliabilities into six unified categories. A wide variety of 911 organic chemicals were selected from the database for QSAR modeling, and 1,504 kinds of different molecular descriptors were calculated, based on their 3D molecular structures as modeled by the Dragon software. Positive (carcinogenic) and negative (non-carcinogenic) chemicals containing various substructures were counted using atom and functional group count descriptors, and the statistical significance of ratios of positives to negatives was tested for those substructures. Very few were judged to be strongly related to carcinogenicity, among substructures known to be responsible for carcinogens as revealed from biomedical studies. In order to develop QSAR models for the prediction of the carcinogenicities of a wide variety of chemicals with a satisfactory performance level, the relationship between the carcinogenicity data with improved reliability and a subset of significant descriptors selected from 1,504 Dragon descriptors was analyzed with a support vector machine (SVM) method: the classification function (SVC) for weighted data in LIBSVM program was used to classify chemicals into two carcinogenic categories (positive or negative), where weights were set depending on the reliabilities of the carcinogenicity data. The quality and stability of the models presented were tested by performing a dual cross–validation procedure. A single SVM model as the first step was developed for all the 911 chemicals using 250 selected descriptors, achieving an overall accuracy level, i.e., positive and negative correct estimate, of about 70%. In order to improve the accuracy of the final model, the 911 chemicals were classified into 20 mutually overlapping subgroups according to contained substructures, a specific SVM model was optimized for each subgroup, and the predicted carcinogenicities of the 911 chemicals were determined by the majorities of the outputs of the corresponding SVM models. The model developed on the basis of grouping of chemicals into 20 substructures predicts the carcinogenicities of a wide variety of chemicals with a satisfactory overall accuracy of approximately 80%.
The ability to assess the toxicity of a chemical substance depends on the available information on the compound and/or its related compounds. Among chemicals currently in commerce, very few are ascertained on their toxicity, and especially reliable data on the carcinogenicity are very limited for pharmaceutical chemicals. Therefore, attempts on the basis of quantitative structure-activity relationship (QSAR) models for estimating the carcinogenicity have been performed. But none of the models so far developed shows satisfactory performance for predicting the carcinogenicity of noncongeneric chemicals from their structures.The support vector machine (SVM) technique was applied to develop a QSAR model that relates the structures of diverse chemicals to their carcinogenicity, and its predictability was compared with that of our previous artificial neural network (ANN) model. The relationship between experimental carcinogenicity data used in the Predictive Toxicology Challenge (PTC) 2000-2001 contest on 454 chemicals and 37 molecular descriptors calculated from their structures alone was analyzed with a software LIBSVM ver.2.85 for support vector regression (SVR). Models were optimized using a cross-validation test for the training dataset, and their performances were evaluated using the test dataset.The training of ANN models took several months using seven PCs to solve the problems such as overtraining, over-fitting and local minima, while SVM gave a just comparable predictability of 74 % with that by ANN, within much shorter computation time. It comes from the advantage of SVM that gives only one global optimum solution after training while ANN gives numerous local minimum solutions. Moreover the prediction accuracy of the SVM model was higher than the best predictability value of 71 % reported in the literature for the same dataset. It is concluded that the support vector machine, a novel nonlinear machine learning approach, leads to a model for predicting the carcinogenicity of noncongeneric chemicals from information on the molecular structure alone with a higher performance than any of the so-far proposed approaches.
近年, 化学物質の毒性に関する情報の取得が地球的な規模で喫緊の課題となっている. しかし, 動物を用いる安全性試験は莫大な時間と費用がかかるため, 毒性が未知の全ての化学物質について動物試験により毒性を評価することは不可能である. また, それらの情報を集録した毒性データベースにも様々な問題点がある. そこで, 構造活性相関, 特にコンピュータを利用した定量的構造活性相関 (QSAR) による毒性予測が化学物質管理の観点から重要になっており, 多くの毒性予測システムが開発されている. しかし, QSARによる既存の毒性予測システムの成績は実用的には不十分であり, 世界中の多くの研究者がこの問題に取り組んでいる. 我々は既存のシステムより高精度かつ高汎用性の予測システムの開発を目指して, ニューラルネットワークを用いた毒性予測手法を研究している. しかし, この問題の解決には信頼度の高い毒性データの収集を始めとして多くの課題が横たわっている.
A three-layered neural network model to predict the hazards of a variety of compounds based on a quantitative structure-activity relationship was developed. The inputs were 10 principal components from 37 kinds of molecular descriptors calculated with MOprograms. For the output the data used in the Predictive Toxicology Challenge (PTC) 2000-2001 contest were employed, containing 454 compounds with the carcinogenic activity of male rats. The total database of 454 compounds was split into training (144 compounds), validation (143) and test (167) sets. To solve the problems such as over-training, over-fitting and local minimum in training the neural network with the error-back-propagation algorithm, various conditions of the network such as the training cycles and neuron numbers of the intermediate layer were optimized. The optimum model showed a correct classification rate close to 74 %, higher than any of the PTC contestants.
Fullerene oxides were the first observed fullerene derivatives and they have naturally attracted attention of both experiment and theory. C60O has represented a long standing case of experiment–theory disagreement, and there has been a similar problem with C60O2—both cases were explained by kinetic rather than thermodynamic control. In this contribution, we report the first computations of C84O. The computations are carried out with the PM3 quantum-chemical semiempirical method. Thermodynamic stabilities are evaluated for the electronic singlet and triplet states. The triplet states are computed at both restricted open-shell Hartree–Fock (ROHF) and unrestricted open-shell Hartree–Fock levels. The computations focus on the two most abundant C84 isomers—D2 and D2d, and especially deal with additions to their shortest and longest bonds. The D2/long structure is the lowest-energy species. The PM3/ROHF energy ordering of the C84O isomers in the first triplet electronic state is exactly the same as in the singlet electronic state and the relative energies are also quite similar. On the other hand, the kinetic stability order is just reversed compared to the thermodynamic order. Hence, for relatively short reaction times the C84O D2/short isomer should primarily be formed (in spite of the fact that it should be, thermodynamically, the least stable in the studied set). The computations point out various stability selection rules controlling production of fullerene-based materials.
In the recycling of poly(vinyl chloride) (PVC), it is required to discriminate every plasticizer for quality control. For this purpose, the near-infrared spectra were measured for 41 kinds of PVC samples with different plasticizers (DINP, DOP, DOA, TOTM and Polyester) and different plasticizer contents (0-49%). A neural-network analysis was applied to the near-infrared spectra pretreated by second-derivative processing. They were discriminated from one another. The neural-network analysis also allowed us to propose a calibration model which predicts the contents of plasticizers in PVC. The correlation coefficient (R) and the root-mean-square error of prediction (RMSEP) for the DINP calibration model were found to be 0.999 and 0.41 wt%, respectively. In comparison, a partial least-squares regression analysis was carried out. The R and RMSEP of the DINP calibration model were calculated to be 0.993 and 1.27 wt%, respectively. It is found that a near-infrared spectra measurement combined with a neural-network analysis is useful for plastic recycling.
高性能リチウムイオン2次電池の開発には,負極炭素材料へのLi/Li(+)の吸蔵・放出機構とその構造の解明が必要である.電子移動に伴うエネルギー損失を最小限にする炭素材料を見い出すために,非経験的分子軌道計算プログラムQ-ChemのRHF/3-21Gを用いて,48種類の炭素骨格に関する電子状態を計算した.その結果,Li/Li(+)のHOMO-LUMOエネルギーの範囲内にある構造を持つためには,tetrabenzo[bc, ef, kl, no]coronene(TBC)を構築ユニットとして数回繰り返した構造が必要であり,特にTBC6が最小骨格であることが明らかになった.
A rapid and intact method has been developed for predicting polyethylene density by near-infrared spectroscopy combined with neural network analysis. Near-infrared spectra in the region of 1.1-2.2 μm wavelength were measured using pellets or powders of twenty-three kinds of polyethylene (PE) with different densities (0.898-0.962 g cm-3). The spectra were used for training a back-propagation neural network after normalized and second-derivative treatments to predict PE density. Although only a small number of spectral data were used for training, a leave-one-out test of neural network analysis has demonstrated good results. In comparison, principal component regression (PCR) analysis and partial least-squares (PLS) regression analysis were applied. The correlation coefficients (R) were calculated to be 1.000, 0.968 and 0.983 for neural network, PCR and PLS analysis, respectively. The root mean square errors of prediction were found to be 0.00026, 0.0043 and 0.0031 g cm-3, respectively. It is found that near-infrared spectroscopy combined with neural network analysis is useful for the efficient and accurate determination of PE density.
構造活性相関により化学物質の構造から有害性を高い精度で予測する手法を開発することを目指して、ニューラルネットワークを用いて発ガン性のデータを解析した。41種類の有機塩素化合物について分子軌道計算などから求まる7種類の記述子を用いてニューラルネットワークを学習し、leave-one-out testを行った結果、的中率93%の予測手法を開発することができた。