Hierarchical Reinforcement Learning (HRL) incorporating intrinsic motivation is a promising approach for sparse reward environments. Stepwise Unified Hierarchical Reinforcement Learning (SUHRL) enables autonomous subgoal generation using Fuzzy ART. However, the conventional SUHRL framework employs the same learning architecture for both the high-level Meta Controller and the low-level Controller, which may limit learning efficiency. This study investigates whether combining different architectures—specifically utilizing Q-Learning for the Meta Controller and Deep Q-Network (DQN) for the Controller—enhances learning efficiency. We conducted comparative experiments using four model variations in a stochastic Grid World environment. The results demonstrate that the hybrid model employing Q-Learning for the high-level policy and DQN for the low-level policy achieved the most efficient learning performance.
The hierarchical reinforcement learning approach incorporates intrinsic motivation into reinforcement learning. In this approach, the agents are hierarchically organized into a higher-level strategy that learns the application order (plan) of sub goals and a lower-level strategy that learns the sequence of actions up to the subgoals. This incorporation has been shown to facilitate problem-solving even in environments with sparse rewards. Stepwise Unified Hierarchical Reinforcement Learning (SUHRL) has been proposed as a hierarchical reinforcement learning algorithm. SUHRL enables autonomous learning without the need for prior knowledge or predefined subgoals. However, learning efficiency is low, and, in some cases, problem solving is not possible. Therefore, we propose the enhancement of SUHRL by constraining the exploration of sub goals based on the number of steps required to transition between them. This improvement is based on Hierarchical Reinforcement Learning with a k-step Adjacency Constraint, a method that can achieve efficient learning when sub goals are predefined. Experimental results using the Frozen Lake environment show that the proposed method exhibits a better learning efficiency than SUHRL.
ABSTRACT Objectives This study aimed to understand the status quo of medical treatments and pregnancy outcomes in patients with Takayasu arteritis (TAK) and children’s birth outcomes. Methods This study retrospectively enrolled patients with TAK who conceived after the disease onset and were managed at medical facilities participating in the Japan Research Committee of the Ministry of Health, Labour, and Welfare for Intractable Vasculitis. Results This study enrolled 51 cases and 68 pregnancies during 2019–21. Of these, 48 cases and 65 pregnancies resulted in delivery and live-born babies. The median age of diagnosis and delivery was 22 and 31 years, respectively. Preconception therapy included prednisolone (PSL) in 51 (78.5%, median 7.5 mg/day), immunosuppressants in 18 (27.7%), and biologics in 12 (18.5%) pregnancies. Six cases underwent surgical treatment before pregnancy. Medications during pregnancy included PSL in 48 (73.8%, median: 9 mg/day), immunosuppressants in 13 (20.0%), and biologics in 9 (13.8%) pregnancies. TAK relapsed in four (6.2%) and eight (12.3%) pregnancies during pregnancy and after delivery, respectively. Additionally, 13/62 (20.9%) preterm infants and 17/59 (28.8%) low-birth-weight infants were observed, and none had serious postnatal abnormalities. Conclusions Most pregnancies in TAK were manageable with PSL at ≤10 mg/day. Relapse during pregnancy and postpartum occurred in <20% of pregnancies.
ABSTRACT Objective To develop a proposal for giant cell arteritis remission criteria in order to implement a treat-to-target algorithm. Methods A task force consisting of 10 rheumatologists, 3 cardiologists, 1 nephrologist, and 1 cardiac surgeon was established in the Large-vessel Vasculitis Group of the Japanese Research Committee of the Ministry of Health, Labour and Welfare for Intractable Vasculitis to conduct a Delphi survey of remission criteria for giant cell arteritis. The survey was circulated among the members over four reiterations with four face-to-face meetings. Items with a mean score of ≥4 were extracted as items for defining remission criteria. Results An initial literature review yielded a total of 117 candidate items for disease activity domains and treatment/comorbidity domains of remission criteria, of which 35 were extracted as disease activity domains (systematic symptoms, signs and symptoms of cranial and large-vessel area, inflammatory markers, and imaging findings). For the treatment/comorbidity domain, ≤5 mg/day of prednisolone 1 year after starting glucocorticoids was extracted. The definition of achievement of remission was the disappearance of active disease in the disease activity domain, normalization of inflammatory markers, and ≤5 mg/day of prednisolone. Conclusion We developed proposals for remission criteria to guide the implementation of a treat-to-target algorithm for giant cell arteritis.
Hierarchical reinforcement learning (HRL) is an approach that incorporates intrinsic motivation mechanisms into reinforcement learning. HRL divides the agent’s internal mechanism into two components: a higher-level policy (application order of subgoals) and a lower-level policy (behavior sequence to subgoals) for problem solving. It has been demonstrated that HRL can solve problems in environments with sparse rewards and environments that require learning of long action sequences, which are difficult to address with conventional reinforcement learning, provided that the definition of subgoals is appropriate. However, existing HRL assumes the availability of predefined subgoals necessary for problem solving and does not provide an algorithm for achieving autonomous reinforcement learning. In this study, we propose stepwise unified hierarchical reinforcement learning (SUHRL), a new reinforcement learning algorithm that introduces a mechanism to gradually generate necessary experiences and appropriate subgoals for problem solving. SUHRL solves problems by stepwise clustering using Fuzzy ART and experience acquisition processing to generate suitable subgoals incrementally. Evaluation experiments conducted on MiniGrid environments and Montezuma’s Revenge demonstrate that the proposed method can generate the required subgoals incrementally and achieve autonomous problem solving.
BackgroundTakayasu arteritis (TAK) is an autoimmune large vessel vasculitis that affects the aorta and its major branches, eventually leading to the development of aortic aneurysm and vascular stenosis or occlusion. This retrospective and prospective study aimed to investigate whether the gut dysbiosis exists in patients with TAK and to identify specific gut microorganisms related to aortic aneurysm formation/progression in TAK.MethodsWe analysed the faecal microbiome of 76 patients with TAK and 56 healthy controls (HCs) using 16S ribosomal RNA sequencing. We examined the relationship between the composition of the gut microbiota and clinical parameters.ResultsThe patients with TAK showed an altered gut microbiota with a higher abundance of oral-derived bacteria, such as Streptococcus and Campylobacter, regardless of the disease activity, than HCs. This increase was significantly associated with the administration of a proton pump inhibitor used for preventing gastric ulcers in patients treated with aspirin and glucocorticoids. Among patients taking a proton pump inhibitor, Campylobacter was more frequently detected in those who underwent vascular surgeries and endovascular therapy for aortic dilatation than in those who did not. Among the genus of Campylobacter, Campylobacter gracilis in the gut microbiome was significantly associated with clinical events related to aortic aneurysm formation/worsening in patients with TAK. In a prospective analysis, patients with a gut microbiome positive for Campylobacter were significantly more likely to require interventions for aortic dilatation than those who were negative for Campylobacter. Furthermore, patients with TAK who were positive for C. gracilis by polymerase chain reaction showed a tendency to have severe aortic aneurysms.ConclusionsA specific increase in oral-derived Campylobacter in the gut may be a novel predictor of aortic aneurysm formation/progression in patients with TAK.
Background: Pulmonary arterial hypertension (PAH) is a type of pulmonary hypertension (PH) characterized by obliterative pulmonary vascular remodeling, resulting in right-sided heart failure. Although the pathogenesis of PAH is not fully understood, inflammatory responses and cytokines have been shown to be associated with PAH, in particular, with connective tissue disease-PAH. In this sense, Regnase-1, an RNase that regulates mRNAs encoding genes related to immune reactions, was investigated in relation to the pathogenesis of PH. Methods: We first examined the expression levels of ZC3H12A (encoding Regnase-1) in peripheral blood mononuclear cells from patients with PH classified under various types of PH, searching for an association between the ZC3H12A expression and clinical features. We then generated mice lacking Regnase-1 in myeloid cells, including alveolar macrophages, and examined right ventricular systolic pressures and histological changes in the lung. We further performed a comprehensive analysis of the transcriptome of alveolar macrophages and pulmonary arteries to identify genes regulated by Regnase-1 in alveolar macrophages. Results: ZC3H12A expression in peripheral blood mononuclear cells was inversely correlated with the prognosis and severity of disease in patients with PH, in particular, in connective tissue disease-PAH. The critical role of Regnase-1 in controlling PAH was also reinforced by the analysis of mice lacking Regnase-1 in alveolar macrophages. These mice spontaneously developed severe PAH, characterized by the elevated right ventricular systolic pressures and irreversible pulmonary vascular remodeling, which recapitulated the pathology of patients with PAH. Transcriptomic analysis of alveolar macrophages and pulmonary arteries of these PAH mice revealed that Il6, Il1b, and Pdgfa/b are potential targets of Regnase-1 in alveolar macrophages in the regulation of PAH. The inhibition of IL-6 (interleukin-6) by an anti–IL-6 receptor antibody or platelet-derived growth factor by imatinib but not IL-1β (interleukin-1β) by anakinra, ameliorated the pathogenesis of PAH. Conclusions: Regnase-1 maintains lung innate immune homeostasis through the control of IL-6 and platelet-derived growth factor in alveolar macrophages, thereby suppressing the development of PAH in mice. Furthermore, the decreased expression of Regnase-1 in various types of PH implies its involvement in PH pathogenesis and may serve as a disease biomarker, and a therapeutic target for PH as well.
Abstract Background Cytosine-phosphate-guanine oligodeoxynucleotide (CpG ODN) (K3)—a novel synthetic single-stranded DNA immune adjuvant for cancer immunotherapy—induces a potential Th1-type immune response against cancer cells. We conducted a phase I study of CpG ODN (K3) in patients with lung cancer to assess its safety and patients’ immune responses. Methods The primary endpoint was the proportion of dose-limiting toxicities (DLTs) at each dose level. Secondary endpoints included safety profile, an immune response, including dynamic changes in immune cell and cytokine production, and progression-free survival (PFS). In a 3 + 3 dose-escalation design, the dosage levels for CpG ODN (K3) were 5 or 10 mg/body via subcutaneous injection and 0.2 mg/kg via intravenous administration on days 1, 8, 15, and 29. Results Nine patients (eight non-small-cell lung cancer; one small-cell lung cancer) were enrolled. We found no DLTs at any dose level and observed no serious treatment-related adverse events. The median observation period after registration was 55 days (range: 46–181 days). Serum IFN-α2 levels, but not inflammatory cytokines, increased in six patients after the third administration of CpG ODN (K3) (mean value: from 2.67 pg/mL to 3.61 pg/mL after 24 hours). Serum IFN-γ (mean value, from 9.07 pg/mL to 12.7 pg/m after 24 hours) and CXCL10 levels (mean value, from 351 pg/mL to 676 pg/mL after 24 hours) also increased in eight patients after the third administration. During the treatment course, the percentage of T-bet-expressing CD8+ T cells gradually increased (mean, 49.8% at baseline and 59.1% at day 29, p = 0.0273). Interestingly, both T-bet-expressing effector memory (mean, 52.7% at baseline and 63.7% at day 29, p = 0.0195) and terminally differentiated effector memory (mean, 82.3% at baseline and 90.0% at day 29, p = 0.0039) CD8+ T cells significantly increased. The median PFS was 398 days. Conclusions This is the first clinical study showing that CpG ODN (K3) activated innate immunity and elicited Th1-type adaptive immune response and cytotoxic activity in cancer patients. CpG ODN (K3) was well tolerated at the dose settings tested, although the maximum tolerated dose was not determined. Trial registration UMIN-CTR number 000023276. Registered 1 September 2016, https://upload.umin.ac.jp/cgi-open-bin/ctr/ctr_view.cgi?recptno=R000026649
Pulmonary arterial hypertension (PAH) is a devastating disease characterized by arteriopathy in the small to medium-sized distal pulmonary arteries, often accompanied by infiltration of inflammatory cells. Aryl hydrocarbon receptor (AHR), a nuclear receptor/transcription factor, detoxifies xenobiotics and regulates the differentiation and function of various immune cells. However, the role of AHR in the pathogenesis of PAH is largely unknown. Here, we explore the role of AHR in the pathogenesis of PAH. AHR agonistic activity in serum was significantly higher in PAH patients than in healthy volunteers and was associated with poor prognosis of PAH. Sprague-Dawley rats treated with the potent endogenous AHR agonist, 6-formylindolo[3,2-b]carbazole, in combination with hypoxia develop severe pulmonary hypertension (PH) with plexiform-like lesions, whereas Sprague-Dawley rats treated with the potent vascular endothelial growth factor receptor 2 inhibitors did not. Ahr-knockout (Ahr-/- ) rats generated using the CRISPR/Cas9 system did not develop PH in the SU5416/hypoxia model. A diet containing Qing-Dai, a Chinese herbal drug, in combination with hypoxia led to development of PH in Ahr+/+ rats, but not in Ahr-/- rats. RNA-seq analysis, chromatin immunoprecipitation (ChIP)-seq analysis, immunohistochemical analysis, and bone marrow transplantation experiments show that activation of several inflammatory signaling pathways was up-regulated in endothelial cells and peripheral blood mononuclear cells, which led to infiltration of CD4+ IL-21+ T cells and MRC1+ macrophages into vascular lesions in an AHR-dependent manner. Taken together, AHR plays crucial roles in the development and progression of PAH, and the AHR-signaling pathway represents a promising therapeutic target for PAH.
BACKGROUNDPulmonary arterial hypertension (PAH), particularly connective tissue disease-associated PAH (CTD-PAH), is a progressive disease and novel therapeutic agents based on the specific molecular pathogenesis are desired. In the pathogenesis of CTD-PAH, inflammation, immune cell abnormality, and fibrosis play important roles. However, the existing mouse pulmonary hypertension (PH) models do not reflect these features enough. The relationship between inflammation and hypoxia is still unclear.Methods and Results:Intraperitoneal administration of pristane, a kind of mineral oil, and exposure to chronic hypoxia were combined, and this model is referred to as pristane/hypoxia (PriHx) mice. Hemodynamic and histological analyses showed that the PriHx mice showed a more severe phenotype of PH than pristane or hypoxia alone. Immunohistological and flow cytometric analyses revealed infiltration of immune cells, including hemosiderin-laden macrophages and activated CD4+helper T lymphocytes in the lungs of PriHx mice. Pristane administration exacerbated lung fibrosis and elevated the expression of fibrosis-related genes. Inflammation-related genes such asIl6andCxcl2were also upregulated in the lungs of PriHx mice, and interleukin (IL)-6 blockade by monoclonal anti-IL-6 receptor antibody MR16-1 ameliorated PH of PriHx mice.CONCLUSIONSA PriHx model, a novel mouse model of PH reflecting the pathological features of CTD-PAH, was developed through a combination of pristane administration and exposure to chronic hypoxia.
TAFRO syndrome is a newly proposed disease that is characterised by thrombocytopenia, anasarca, fever, reticulin fibrosis (or renal dysfunction), and organomegaly. Generally, high doses of corticosteroids are recommended for the initial treatment of TAFRO syndrome; however, some patients experience prolonged refractory thrombocytopenia after initiating such therapies. If corticosteroid treatment alone is ineffective, additional immunosuppressive therapies such as cyclosporine A are recommended. Since long-term use of immunosuppressive therapies with TAFRO syndrome sometimes causes serious infection, it is important to recognise the time to recovery from thrombocytopenia. In this study, we investigated how long it took to recover from thrombocytopenia, to aid clinicians in decision-making regarding the need to strengthen treatment for prolonged thrombocytopenia. Here, we describe three of our patients with TAFRO syndrome exhibiting prolonged thrombocytopenia. We also investigated the median period to recovery from this complication (defined as the time to increase the platelet count above 50,000/µL) after the initiation of high-dose corticosteroid treatment in our 3 cases and 38 peer-reviewed cases. We found that it took our patients 61 days to recover from thrombocytopenia; in comparison, our investigation of the 38 peer-reviewed case reports revealed a median recovery time of 47.5 days among previously reported patients. We showed the time to recovery from thrombocytopenia in patients with TAFRO syndrome for the first time. Our findings ought to be useful for decision-making among clinicians regarding the administration of other immunosuppressive treatments in addition to corticosteroid.
Recently, the number of households comprising only elderly people(60 years old or older) has increased because of the falling birth rate and the aging population. According to a recent Japanese Statistics Bureau report, the total population was estimated to be 126.59 million among which 35.22 million people were elderly. Furthermore, the Ministry of Health, Labor, and Welfare predicted a shortage of approximately 380,000 nursing care staff in Japan by 2025 [1], which is the year in which the baby-boomer generation is expected to become more than 75 years old. As the number of users of nursing care services increases, 2.53 million nursing staff will become necessary by 2025; however, it is expected that only 2.15 million staff will be present based on the current rate of increase. According to the official release of the sufficiency rate associated with the number of nursing care staff actually required to serve the number of people who requires them, which increase with the aging population, there will be a shortage of care workers of approximately 200,000 in 2020 and of approximately 380,000 in 2025. Therefore, we have developed a video-surveillance system capable of detecting an elderly person falling in the absence of care workers.
Context is defined as information that represents the environment surrounding us. A system to provide services to users using context is called a context-aware system (CAS), and it has been widely studied. However, many conventional CASs are not flexible. As conventional CASs cannot be freely customized by end-users, they lack adaptability to environments, making them difficult to use in various environments. To solve this problem, we propose a user-oriented context-aware system architecture (U-CASA) for building ad-hoc smart rooms. This architecture allows end-users to install various IoT sensor devices to the target environment and to make it a smart room. The end-user can also freely customize the system settings and can link various contexts to a variety of services. In order to evaluate the proposed architecture, we built two ad-hoc smart rooms based on U-CASA. As a result of the demonstration experiment with behavior scenarios, we confirmed that two smart rooms can perform properly and can provide services despite their different environments.
The scene recognition is one of the most important tasks for estimating ambient attributes from the image data. The ambient attributes are scene traits for specifying the meanings of human activities. We consider that the estimation of ambient attributes is required for understanding human activities in various scenes. Thus this paper proposes a novel method of the scene recognition by the object detector. The proposed method estimates a scene using a histogram of objects, which is called Bag of Objects, based on the result of object detection. To evaluate the proposed method, a simulation experiment has been done with a lot of image data we collected from the web. As the result, we show that the average of scene recognition accuracy is 0.58 for 26 scene categories.
In this paper, we propose an agent-based architecture for remote collaboration support systems that enables the exchange of synchronous and asynchronous multimedia streams at remote sites using an Internet of Multimedia Things (IoMT) approach. First, we design and implement Internet of Things (IoT) applications that contain simple sensors and actuators. These applications are modularized into agent based subsystems that can be incorporated into an IoT application with agent operations. Our aim is to develop remote collaboration support applications composed of video, data, and document channels. Because the applications will exchange enormous multimedia streams between remote sites, we propose a novel IoMT system architecture composed of several channel types that consist of various resource and network components. Users can dynamically incorporate these channels into applications to update the IoMT system's effects. Finally, we demonstrate and discuss the experimental results of our application to validate its ability to rapidly supply multimedia resources via multi-agent collaboration. Adding to its novelty, the developed system is in practical use for collaboration between France and Japan.
Software-defined networking (SDN) not only lessens the burden of network operators but also accelerates the provision of high-quality services to various users by speeding up changes to the network configuration. Recently, Wireless Networking (WN) has gained importance as a social infrastructure to connect a large number of devices and cloud applications. Software definition technology for wireless networks is referred to as software-defined wireless networking (SDWN). We hereby propose a cognitive software-defined wireless networking (CS-DWN) to enhance SDWN by applying a cognitive cycle to SDWN in addition to the development of a prototype to validate this concept.
The new technology to integrate Edge Computing and Cloud Computing is called Internet of Things (IoT).Current research on IoT mainly focuses on how to connect Edge Computing to Cloud Computing.However, Oihui Wu, et al. argued that simply connecting them is not enough; beyond that, objects should have the capability to learn, think, and understand both physical and social worlds by themselves.An architecture of IoT applications is being developed to connect Agent Spaces, including Personal Assistant Agents and Service Agents in order to support user activity at the office, to both the Edge and the Cloud.In this paper, an architecture of IoT applications is proposed connecting Edge Resources and Cloud Resources.According to the proposed architecture, an Agent Space is designed in which agents can control the resources via programs named Resource Connectors.Finally, an IoT application was prototyped to support users working at the office with a Personal Assistant Agent in an Agent Space that has a vocal interface with users and a web interface.
This paper examines the influence of ramp-event of aggregated power output fluctuation of photovoltaic power generation system (PVS) on the power system frequency. A numerical simulation model of economic load dispatching control (EDC) and load - frequency control (LFC) is utilized together with a PVS power output forecasting model and a unit commitment (UC) scheduling model developed in our preceding study. In the case of ramp event with long duration and rapid ramp rate, the power output of controllable generators with high load following capability reaches to the limit even though the power output change of low load following capability generators is still available. As a result, the frequency deviation from the acceptable range occurs. When the load dispatch scheme is temporary switched from the conventional EDC using an equal incremental fuel cost rule to the dispatching based on the capacity without the consideration of fuel cost, the aggregated load-following capability can be preserved and the frequency deviation can be avoided.