Para el 2050, cerca del 52% de la población mundial vivirá en zonas con escasez hídrica (The Economist, 2019), lo cual, sin duda, refleja que la crisis global del agua es uno de los principales desafíos de la humanidad en el siglo XXI. Sin embargo, la crisis del agua no es solo un asunto de escasez del recurso hídrico, sino, ante todo, una problemática vinculada con su gobernanza y, precisamente, uno de los escenarios en donde está más presente dicha crisis son las cuencas transfronterizas. En la actualidad, en el mundo hay 150 países que comparten 310 cuencas transfronterizas, las cuales abarcan el 47.1% del área terrestre del planeta. (McCracken y Wolf, 2019). La problemática de gobernabilidad de las cuencas transfronterizas es alimentada por la evidente tensión entre el orden westfaliano de Estados-nación y la naturaleza física del agua. En otras palabras, tenemos un recurso que se mueve atravesando fronteras y, al mismo tiempo, los límites de los Estados nacionales, cuyas fronteras han sido pensadas y definidas como elementos estáticos, tal dicotomía es ejemplificada con los casos de las cuencas de los ríos Nilo, Amazonas y Mekong. En conclusión, el principal desafío que enfrenta la gestión de cuencas transfronterizas es la construcción de instituciones que logren, por un lado, superar la lógica de administración territorial basada en el enfoque de soberanía nacional y, por otro, que incorpore la noción de derechos y deberes compartidos dentro de nuevos marcos institucionales de cooperación transfronteriza.
Objective: Axillary management in elderly patients with early breast cancer and clinically negative axillary nodes is controversial. This study aimed to evaluate clinico-histopathological and survival data in breast cancer patients aged 80 years or older with a negative axillary clinical and ultrasound examination not undergoing axillary lymph node investigation Patients and Method: A retrospective study of 36 patients who met the following inclusion criteria: aged 80 years or more, with a diagnosis of early breast cancer and clinically and ultrasound node negative breast cancer, and in a state of complete physical, mental well-being. The patients were treated surgically without axillary dissection. Clinico-histopathological, treatment and survival characteristics were evaluated. Results: A total of 36 patients were studied with an average age of 83.5 years (range, 80–91 years), and a median follow-up of 39.7 months (range: 0.5–168 months). Of these, 1 patient had bilateral breast cancer. Most patients were treated with partial mastectomy (64.9%). Almost all of the patients received adjuvant hormonal therapy (91.7%), almost one third received adjuvant radiotherapy (30.6%). Infiltrating ductal carcinoma represented 62% of the total. Average tumor size of 17.5 mm (range, 2– 50 mm). The most frequent molecular subtype was luminal A (54.1%). 8.3% of the patients had a relapse, none in axilla. Disease free survival was found to be a mean of 141.7 months, while five-year disease-free survival rate was 80.9%. The overall survival average time was 95.5months, while the five-year overall survival rate was 57%.Table 1:Clinical and Histopathological characteristicsAge (years) N = 36 (1 bilateral)80–8424 (66.7)85–9112 (33.3)Media (STD)83.5 (0.49)Tumor characteristicsTumor size (mm) N = 37≤2029 (78.4)<20 ≤ 508 (21.6)Media (STD) (range)17.5(1.52) (5–50)Histological type N = 37IDC23 (62.2)ILC4 (10.8)Others10 (27)Molecular Subtypes N = 37Luminal A20 (54.1)Luminal B10 (27)Luminal B-H5 (13.5)HER21 (2.7)Triple negative1 (2.7)Type of Surgery N = 37BCS24 (64.9)TM£1 patient with bilateral mastectomy13 (35.1)Treatment N = 36Adjuvant Hormonotherapy33 (91.7)Adjuvant Radiotherapy11 (30.6)Neoadjuvant Hormonotherapy9 (25)IDC: infiltrating ductal carcinoma, ILC: infiltrating lobullar carcinoma, Others: muninous carcinoma, papillary carcinoma, cribiform carcinoma, ductal carcinoma in situ. BCS: breast conservative surgery. TM: Total mastectomy.£ 1 patient with bilateral mastectomy Open table in a new tab IDC: infiltrating ductal carcinoma, ILC: infiltrating lobullar carcinoma, Others: muninous carcinoma, papillary carcinoma, cribiform carcinoma, ductal carcinoma in situ. BCS: breast conservative surgery. TM: Total mastectomy. Conclusion: Results from the present investigation suggest that axillary lymph node biopsy could be omitted in women 80 years or older with clinically node negative breast cancer treated with breast surgery and adjuvant therapy. No conflict of interest.
Background: Neoadjuvant therapy (NAT) increases the probability of breast-conserving surgery, establishing NAT as a viable option in patients who are not eligible for breast-conserving surgery when first diagnosed. NAT also allows an increase in rates of complete pathological response (pCR). We analyzed two groups of patients with breast cancer treated with NAT, one group showed pCRversus a second group that showed partial pathological response (pPR). We determined clinicopathological characteristics associated with pCR, as well as its role as a prognostic factor for disease-free survival (DFS).
Video sharing (e.g., YouTube, Vimeo, Facebook, TikTok) accounts for the majority of internet traffic, and video processing is also foundational to several other key workloads (video conferencing, virtual/augmented reality, cloud gaming, video in Internet-of-Things devices, etc.). The importance of these workloads motivates larger video processing infrastructures and – with the slowing of Moore’s law – specialized hardware accelerators to deliver more computing at higher efficiencies. This paper describes the design and deployment, at scale, of a new accelerator targeted at warehouse-scale video transcoding. We present our hardware design including a new accelerator building block – the video coding unit (VCU) – and discuss key design trade-offs for balanced systems at data center scale and co-designing accelerators with large-scale distributed software systems. We evaluate these accelerators “in the wild" serving live data center jobs, demonstrating 20-33x improved efficiency over our prior well-tuned non-accelerated baseline. Our design also enables effective adaptation to changing bottlenecks and improved failure management, and new workload capabilities not otherwise possible with prior systems. To the best of our knowledge, this is the first work to discuss video acceleration at scale in large warehouse-scale environments.
The fundamental understanding and tailoring of material properties play a fundamental role on the performance of tin (II) sulfide (SnS)-based photovoltaic devices. In this regard, this work reports on the deposition of single-phase, p-type SnS thin films synthesized by close spaced vapor transport (CSVT) and the impact of employing argon (Ar) and bare vacuum (air) atmospheres on CSVT-SnS thin film properties is presented and compared for the first time. The analysis of film properties was performed by structural , directional , morphological , topographical , along with thermoelectric and optoelectronic characterizations of the samples. It is demonstrated that by changing the CSVT annealing atmosphere, the SnS grains change their orientation from perpendicular (respect to substrate) in Ar to parallel in a bare vacuumed (air) conditions. Furthermore, Hall measurements show that the hole concentration are mostly the same but that hole mobility depends, in turn, on the direction of preferred orientation of the SnS crystals. This translates into a twice higher lateral conductivity on the airrespect to the argon-SnS thin films. In this way, the different atmospheric exhibit a profound impact on the physical , morphological and structurally orientated properties of the SnS grains. These results demonstrate that photovoltaic grade p-type SnS films can be deposited in air without the need of an inert gas atmosphere which opens the way to a cost reduction in the fabrication of SnS-based photovoltaic devices.
This paper presents vbench, a publicly available benchmark for cloud video services. We are the first study, to the best of our knowledge, to characterize the emerging video-as-a-service workload. Unlike prior video processing benchmarks, vbench's videos are algorithmically selected to represent a large commercial corpus of millions of videos. Reflecting the complex infrastructure that processes and hosts these videos, vbench includes carefully constructed metrics and baselines. The combination of validated corpus, baselines, and metrics reveal nuanced tradeoffs between speed, quality, and compression. We demonstrate the importance of video selection with a microarchitectural study of cache, branch, and SIMD behavior. vbench reveals trends from the commercial corpus that are not visible in other video corpuses. Our experiments with GPUs under vbench's scoring scenarios reveal that context is critical: GPUs are well suited for live-streaming, while for video-on-demand shift costs from compute to storage and network. Counterintuitively, they are not viable for popular videos, for which highly compressed, high quality copies are required. We instead find that popular videos are currently well-served by the current trajectory of software encoders.
This paper presents vbench, a publicly available benchmark for cloud video services. We are the first study, to the best of our knowledge, to characterize the emerging video-as-a-service workload. Unlike prior video processing benchmarks, vbench's videos are algorithmically selected to represent a large commercial corpus of millions of videos. Reflecting the complex infrastructure that processes and hosts these videos, vbench includes carefully constructed metrics and baselines. The combination of validated corpus, baselines, and metrics reveal nuanced tradeoffs between speed, quality, and compression. We demonstrate the importance of video selection with a microarchitectural study of cache, branch, and SIMD behavior. vbench reveals trends from the commercial corpus that are not visible in other video corpuses. Our experiments with GPUs under vbench's scoring scenarios reveal that context is critical: GPUs are well suited for live-streaming, while for video-on-demand shift costs from compute to storage and network. Counterintuitively, they are not viable for popular videos, for which highly compressed, high quality copies are required. We instead find that popular videos are currently well-served by the current trajectory of software encoders.
The role of monoamines in epilepsy is not clearly reported in the literature. Seizures are associated with the release of catecholamines when the cerebrospinal fluid (CSF) is analyzed in a less than 2 hours period after the epileptic event. However, no detailed studies regarding intercritical alterations have been described. Clinical data and levels of neurotransmitters (biogenic amines) in CSF of 90 patients with early epileptic encephalopathies followed in HSJD, Barcelona, were recruited. Data about the epileptic syndrome type and other neurological features, electroencephalography, genetic studies, brain magnetic resonance imaging, and extensive metabolic screening were collected. The series was composed by 51 females and 39 males with a median age of 1.37 years at the moment of lumbar puncture and 0.52 years at the epilepsy onset. Thirty-one patients had abnormal levels of biogenic amines (34.4%). Sixteen patients had isolated alteration of 5-HIAA (serotonin metabolite) and 10 abnormal isolated HVA levels (dopamine metabolite); five patients had a combined HVA+5-HIAA decrease. Twenty-five patients had a positive genetic diagnosis. Onset age was the only factor related to higher probability of NT depletion. So far only four patients with low CSF NT levels have been treated (one with 5-hidroxytriptophan and three with combined L-dopa+carbidopa and 5-hydroxytriptophan). All of them showed a sustained reduction in seizures (very striking in two patients and moderate in the other two) and improvement in other neurodevelopmental skills. Although this is an ongoing study and requires further analysis, biogenic amines seem to be importantly affected in EE, in particular in very young children. Studies about therapeutic replacement in long series of patients are badly needed to establish formal treatment recommendations, but these preliminary results are promising.
Memory stalls are a significant source of performance degradation in modern processors. Data prefetching is a widely adopted and well studied technique used to alleviate this problem. Prefetching can be performed by the hardware, or be initiated and controlled by software. Among software controlled prefetching we find a wide variety of schemes, including runtime-directed prefetching and more specifically runtime-directed block prefetching. This paper proposes a hybrid prefetching mechanism that integrates a software driven block prefetcher with existing hardware prefetching techniques. Our runtime-assisted software prefetcher brings large blocks of data on-chip with the support of a low cost hardware engine, and synergizes with existing hardware prefetchers that manage locality at a finer granularity. The runtime system that drives the prefetch engine dynamically selects which cache to prefetch to. Our evaluation on a set of scientific benchmarks obtains a maximum speed up of 32 and 10 % on average compared to a baseline with hardware prefetching only. As a result, we also achieve a reduction of up to 18 and 3 % on average in energy-to-solution.
This work reports results of a study carried out to improve the optical, electrical and structural properties of CZTS films grown by spray pyrolysis in a one-step process using a precursor solution prepared dissolving thiourea and salts of Cu, Sn and Zn in a solvent constituted by a mixture of dimethyl sulfoxide (DMSO) and acetone. The improvement of the properties of the CZTS films was achieved through a parameters study performed by using a factorial experimental design 2 3 face centered with reply in the central point. Special emphasis was done on studying the influence of substrate temperature (T s ), carrier gas pressure (P g ) and spray pulse time (tsp) on the optical, electrical and structural properties of the CZTS films. The study revealed that the t SD and T s as well as their interaction are the parameters that most critically affect the above mentioned properties.
GPUs achieve high throughput and power efficiency by employing many small single instruction multiple thread (SIMT) cores. To minimize scheduling logic and performance variance they utilize a uniform memory system and leverage strong data parallelism exposed via the programming model. With Moore’s law slowing, for GPUs to continue scaling performance (which largely depends on SIMT core count) they are likely to embrace multi-socket designs where transistors are more readily available. However when moving to such designs, maintaining the illusion of a uniform memory system is increasingly difficult. In this work we investigate multi-socket non-uniform memory access (NUMA) GPU designs and show that significant changes are needed to both the GPU interconnect and cache architectures to achieve performance scalability. We show that application phase effects can be exploited allowing GPU sockets to dynamically optimize their individual interconnect and cache policies, minimizing the impact of NUMA effects. Our NUMA-aware GPU outperforms a single GPU by $1.5 \times, 2.3 \times$, and $3.2 \times$ while achieving 89%, 84%, and 76% of theoretical application scalability in 2, 4, and 8 sockets designs respectively. Implementable today, NUMA-aware multi-socket GPUs may be a promising candidate for scaling GPU performance beyond a single socket.CCS CONCEPTS• Computing methodologies → Graphics processors; • Computer systems organization → Single instruction, multiple data;
High performance computing (HPC) applications have parallel code sections that must scale to large numbers of cores, which makes them sensitive to serial regions. Current supercomputing systems with heterogeneous or asymmetric CMPs (ACMP) combine few high-performance big cores for serial regions, together with many low-power lean cores for throughput computing. The low requirements of HPC applications in the core front-end lead some designs, such as SMT and GPU cores, to share front-end structures including the instruction cache (I-cache). However, little work exists to analyze the benefit of sharing the I-cache among full cores, which seems compelling as a solution to reduce silicon area and power. This paper analyzes the performance, power and area impact of such a design on an ACMP with one high-performance core and multiple low-power cores. Having identified that multiple cores run the same code during parallel regions, the lean cores share the I-cache with the intent of benefiting from mutual prefetching, without increasing the average access latency. Our exploration of the multiple parameters finds the sweet spot on a wide interconnect to access the shared I-cache and the inclusion of a few line buffers to provide the required bandwidth and latency to sustain performance. The projections with McPAT and a rich set of HPC benchmarks show 11% area savings with a 5% energy reduction at no performance cost.
There is a need to increase performance under the same power and area envelope to achieve Exascale technology in high performance computing (HPC). The today's chip multiprocessor (CMP) design is tailored by traditional desktop and server workloads, different from parallel applications commonly run in HPC. In this work, we focus on the HPC code characteristics and processor front-end which factors around 30% of core power and area on the emerging lean-core type of processors used in HPC. Separating serial from parallel code sections inside applications, we characterize three HPC benchmark suites and compare them to a traditional set of desktop integer workloads. HPC applications have biased and mostly backward taken branches, small dynamic instruction footprints, and long basic blocks. Our findings suggest smaller branch predictors (BP) with the additional loop BP, smaller branch target buffers (BTB), and smaller L1 instruction caches (I-cache) with wider lines. Still, the aforementioned downsizing applies only to the cores meant to run parallel code. The difference between serial and parallel code sections in HPC applications points to an asymmetric CMP design, with one baseline core for sequential and many HPCtailored cores designed for parallel code. Predictions using Sniper simulator and McPAT show that an HPC-tailored lean core saves 16% of the core area and 7% of power compared to a baseline core, without performance loss. Using the area savings to add an extra core, an asymmetric CMP with one baseline and eight tailored cores has the same area budget as a symmetric CMP composed out of eight baseline cores demanding 4% more power and providing 12% shorter execution time on average.
The European Autism Information System project highlighted the lack of systematic and reliable data relating to the prevalence of autism spectrum disorders in Europe. A protocol for the study of ASD prevalence at European level was developed to facilitate a common format for screening and diagnosing children across the EU. This is the first study to operationalise and screen national school children for ASDs using this protocol. National school children 6–11 years (N = 7951) were screened males 54 % (N = 4268) females 46 % (N = 3683). Screening children for ASD implementing the EAIS protocol using the Social Communication Questionnaire (Rutter et al. in Social Communication Questionnaire (SCQ). Western Psychological Services, Los Angeles, ) as a first level screening instrument in a non-clinical setting of Irish national schools was demonstrated.
High-performance computing (HPC) is recognized as one of the pillars for further progress in science, industry, medicine, and education. Current HPC systems are being developed to overcome emerging architectural challenges in order to reach Exascale level of performance, projected for the year 2020. The much larger embedded and mobile market allows for rapid development of intellectual property (IP) blocks and provides more flexibility in designing an application-specific system-on-chip (SoC), in turn providing the possibility in balancing performance, energy-efficiency, and cost. In the Mont-Blanc project, we advocate for HPC systems being built from such commodity IP blocks, currently used in embedded and mobile SoCs. As a first demonstrator of such an approach, we present the Mont-Blanc prototype; the first HPC system built with commodity SoCs, memories, and network interface cards (NICs) from the embedded and mobile domain, and off-the-shelf HPC networking, storage, cooling, and integration solutions. We present the system's architecture and evaluate both performance and energy efficiency. Further, we compare the system's abilities against a production level supercomputer. At the end, we discuss parallel scalability and estimate the maximum scalability point of this approach across a set of applications.
......Recent packaging technologies that enable DRAM chips to be stacked inside the processor package or on top of the processor chip can lower DRAM energy-per-bit costs, provide wider interfaces, and deliver substantially higher memory bandwidth. However, these technologies are limited in capacity and come at a higher price than traditional off-package memories, so system designers must balance price, performance, and capacity tradeoffs. The most obvious means to achieve this balance is to employ both onand off-package memory in a heterogeneous memory architecture. However, designers must then decide whether to deploy the on-package memory as an additional cache hierarchy level (controlled by hardware or software) or as a memory peer to the offpackage DRAM in a nonuniform memory access (NUMA) configuration. Figure 1 shows a generic memory hierarchy with increasing capacities, access latencies, and energy characteristics. Bandwidth and cost per bit decrease with the increase in distance from the computing core. Table 1 summarizes a memory hierarchy actual bandwidth, latency, and capacity values for an example quad-core CPU at 3 GHz. The gap in bandwidth and capacity between on-chip last-level caches and external DRAMs is usually higher than 10 . Tiers in the memory hierarchy can be explicitly managed by software (for example, CPU registers managed by the compiler or file-system buffers managed by the OS) or transparently managed by hardware (for example, static RAM [SRAM] caches). Hardware caches provide a graceful mechanism to improve latency and bandwidth without adding complexity in the software needed to detect and exploit locality. However, hardware caches introduce area and energy overheads due to tag and data management. With the introduction of onpackage memories, the question must be reevaluated, yet again, to determine if the memory hierarchy has sufficient room for another tier of caching between these new heterogeneous memory technologies. This article presents a model and analysis of energy, bandwidth, and latency for current Evgeny Bolotin David Nellans Oreste Villa Mike O’Connor Alex Ramirez Stephen W. Keckler
Nacho Navarro合作论文数Departament Arquitectura Computadors8