
Esophageal squamous cell carcinoma (ESCC) remains a major cause of cancer-related mortality, and prognosis depends strongly on detection at a curable stage. Endoscopy is central to screening and diagnosis, but subtle flat lesions, operator dependence, cognitive fatigue, lesion-location blind spots, and variability in interpretation contribute to missed or delayed diagnosis. Artificial intelligence (AI), particularly deep learning applied to white-light imaging, narrow-band imaging, blue-light imaging, magnifying endoscopy, Lugol chromoendoscopy, video endoscopy, and endocytoscopy, has shown clinically meaningful potential for lesion detection, margin delineation, invasion-depth estimation, and microvascular-pattern classification. Following peer-review feedback, this article has been reframed as a structured mini-review and evidence appraisal rather than a formal systematic review or meta-analysis. We summarize representative studies published mainly from 2019 to 2026, describe the literature-identification scope and eligibility criteria, and critically appraise the evidence using domains adapted from diagnostic-accuracy, prediction-model, and medical-imaging AI reporting frameworks. In this review, multimodal AI refers specifically to complementary endoscopic inputs rather than routine integration of histopathological or genomic data into deployed ESCC endoscopic systems. Current evidence suggests that many AI systems achieve high sensitivity in enriched image datasets, and several video-based, prospective, or randomized studies support translational feasibility. However, major limitations remain: most studies are retrospective; many use single-center or high-quality still-image datasets; confidence intervals, calibration, uncertainty quantification, subgroup analysis, failure-mode reporting, and latency benchmarks are inconsistent; and external validation across devices, operators, patient spectra, and live-video workflows remains limited.
Online voice-based applications and speech communication have grown as a result of the revolutionary rise of smart gadgets and social media. The rapid advancement of deep learning (DL) has transformed the field of audio processing, enabling smooth human-computer interaction. DL approaches have been used to develop speech-to-text (STT) systems across various languages and topics. These models require a large amount of training data: extensive corpora of continuous speech utterances collected from numerous speakers, along with their corresponding transcripts. In this study, we explore the use of state-of-the-art pre-trained models like Wav2Vec2 XLSR-53 and Whisper-small for developing STT systems in the Telugu language, addressing the challenge of limited data availability and demonstrating satisfactory results [23.62% Word Error Rate (WER), 4.12% Character Error Rate (CER)] even when fine-tuned on a smaller dataset. To evaluate model performance, we employed a k-fold cross-validation approach with values of k = 2 to k = 5, and compared the results with the conventional train-test split method. The results indicate that at k = 5, the Wav2Vec2 XLSR-53 model achieved a cross-validation Character Error Rate(WER) of 23.62% and a Character Error Rate (CER) of 4.12%, and the Whisper-small model yielded a cross-validation WER of 28.67% and a CER of 5.48%. These findings suggest that the k-fold cross-validation strategy, at k = 5, enhances the robustness of Wav2Vec2 XLSR-53 in low-resource language scenarios where training data is limited. For Whisper-small, however, the baseline train-test split (WER: 27.76%) outperformed all k-fold configurations tested (k = 2: 32.05%, k = 3: 29.72%, k = 4: 29.40%, k = 5: 28.67%), indicating that the benefit of k-fold cross-validation observed for Wav2Vec2 XLSR-53 does not generalize across model architectures. Additionally, the models were evaluated on unseen noisy data, and both models demonstrated satisfactory performance.
Modern clinical epidemiology and artificial intelligence are increasingly driven by an idealized premise: the belief that massive databases and advanced machine learning can redeem causal inference from observational uncertainty. Behind the facade of multi-million-record cohorts, however, a profound epistemological crisis festers, revealing a potentially sick artificial intelligence. Modern analytical systems leverage massive sample sizes to generate ultra-significant asymptotic p-values that reflect mathematical outcomes rather than biological reality. This article diagnoses this systemic pathology, the uncritical application of Gaussian asymptotic statistics to sparse, discrete, and rare medical event counts governed by Poisson distributions. To cure this condition, we introduce a methodological antidote: an inference stress test that repurposes the exact lower boundary of the exact Poisson confidence interval as a formal null benchmark, utilizing an event-anchored standard error. Across three clinical case studies, spanning autoimmune dermatology baselines, expanded matching cohorts, and oncological lifestyle investigations, our stress test provides evidence that nominal asymptotic significance may sometimes fail to clear the structural resistance threshold, revealing the underlying fragility of scale-induced signals. This proposed framework serves as an essential epistemological filter, separating genuine biological discovery from digital outcomes and ensuring that future medical AI systems learn only from structurally validated knowledge.
Automatic cardiac arrhythmia detection using electrocardiogram (ECG) signals is essential for early diagnosis of cardiovascular conditions. Most of the previously existing systems for arrhythmia detection depend on fully supervised learning approaches. In addition to that, since the ECG patient data is centrally stored, it can raise concerns like privacy and security in healthcare environments. To address these issues, this work explores a semi-supervised federated learning framework that detects arrhythmia using ECG signals. The MIT-BIH Arrhythmia dataset is used from which the heartbeat segments are extracted and relevant features are obtained. The framework incorporates the use of multi-level feature extraction, adaptive feature fusion and pseudo-labeling to utilize both labeled and unlabeled data efficiently. In order to improve the performance of the model, a weighted federated aggregation approach is used where the contribution of each of the clients is considered during the global model updates. Implementation and evaluation of multiple machine learning and deep learning models including Random Forest, Support Vector Machine, K-Nearest Neighbor, AdaBoost, Gradient Boosting and Multi-Layer Perceptron are performed within the proposed framework. This proposed approach achieves an accuracy of 96.81% in semi-supervised federated settings and has an AUC value of 0.9957 that shows high classification reliability. The system helps improve the model learning with limited number of labeled data at the same time ensuring privacy preserving distributed training. The model maintains a competitive performance under federated settings as compared to the centralized training. Among the evaluated models, ensemble and neural network based approaches show superior performance. The framework demonstrates a stable performance across federated training rounds. Overall, the results obtained show the effectiveness of the proposed framework for arrhythmia detection in a simulated federated setting.
Background and aimsArtificial intelligence-based computer-aided detection (CADe) systems have been developed to enhance the adenoma detection rate (ADR) during colonoscopy, but their performance is unknown. We primarily aimed to compare the effectiveness of each CADe system with conventional colonoscopy (CC). As a secondary objective, we performed an exploratory comparison among different CADe systems.MethodsA systematic literature search of 6 databases was conducted to find randomized controlled trials (RCTs) evaluating the use of CADe systems during colonoscopy. A Bayesian network meta-analysis was performed on the included studies using R 4.5.2 software. The primary outcome was ADR.ResultsA total of 21 RCTs involving 19,006 participants were included. Nine CADe systems were compared. For ADR based on modified intention-to-treat (ADR-mITT), DEEP2 (RR 1.37; 95% CrI 1.07–1.76), Eagle-Eye (RR 1.19; 95% CrI 1.02–1.38), and ENDO-AID (RR 1.27; 95% CrI 1.13–1.41) were significantly higher than CC. For ADR based on intention-to-treat (ADR-ITT), AQCS was significantly higher than CC. For adenoma per colonoscopy (APC) and diminutive adenoma detection rate (DADR), ENDO-AID showed significant advantages over CC. However, no significant differences were observed among CADe systems in terms of high-risk lesion detection rate (including advanced adenoma detection rate [AADR] and sessile serrated lesion detection rate [SDR]) or withdrawal time (WT).ConclusionCompared with CC, the CADe system significantly improved ADR. However, given the limited direct comparative evidence and the star-shaped network, the results, especially the differences among systems, should be interpreted cautiously.Systematic review registrationhttps://www.crd.york.ac.uk/PROSPERO/view/CRD420261290518, identifier: CRD420261290518.
Crop recommendation is a vital part of precision agriculture as it helps farmers choose appropriate crops according to the nutrient profile and environmental conditions. This paper presents a crop recommendation framework in which Multi-Layer Perceptron (MLP), XGBoost, and Tab Transformer are first evaluated as baseline prediction models, followed by the proposed Krill Herd Optimization (KHO)-based explainable framework integrated with Explainable Artificial Intelligence (XAI). These eleven parameters are created based on agronomic and environmental aspects, namely: Nitrogen, Phosphorus, Potassium, Copper, Iron, Magnesium, Sulphur, Temperature, Rainfall, pH and Humidity. For better model transparency and to aid informed decision-making, model explanations with SHAP (SHapley Additive Explanations) and LIME (Local Interpretable Model-Agnostic Explanations) are used to identify feature contributions to crop predictions both locally and globally. The experimental results showed that Tab Transformer significantly outperformed the other models, with an accuracy of 0.99, precision of 0.98, recall of 0.99 and F1-score of 0.98. The proposed framework further incorporates Krill Herd Optimization (KHO) to generate optimized nutrient and climate profiles, while Genetic Algorithm (GA), Particle Swarm Optimization (PSO), and Simulated Annealing (SA) are used for comparative evaluation of optimization performance. By combining explainable AI with optimization methods, the framework improves crop suitability prediction and provides transparent insights into the factors influencing crop recommendations, ensuring reliable decision support for practical farming applications. The proposed framework supports precision agriculture by enabling data-driven crop selection, reducing unnecessary fertilizer usage, optimizing crop productivity, and promoting sustainable farming practices.
PurposeWith the growing integration of Artificial Intelligence (AI) into contemporary workplaces, understanding how AI-enabled work systems influence employee wellbeing has become an important concern within HRM research, digital work literature, and future-of-work scholarship.Design/methodology/approachDrawing on the Job Demands–Resources (JD-R) model and Conservation of Resources (COR) theory, this study examines the mediating roles of technostress and psychological resilience in the relationships between AI adoption, digital work intensity, and employee wellbeing, while also testing the moderating role of perceived organisational support. Data were collected from 516 employees in the Indian IT industry and analysed using Partial Least Squares Structural Equation Modelling (PLS-SEM).FindingsThe findings suggest that the direct effect of AI adoption on employee wellbeing remains unclear, while digital work intensity positively influences wellbeing. Both AI adoption and digital work intensity significantly affect technostress and psychological resilience, reflecting the dual nature of AI-enabled work environments. Technostress does not significantly influence employee wellbeing, whereas psychological resilience is a significant predictor and mediator. Perceived organisational support was positively associated with employee wellbeing; however, its moderating effect on the relationship between technostress and employee wellbeing was not significant.Originality/valueThe study is original because it shows that AI-based work systems create both strain and resource-building processes. The results also indicate that psychological resilience has a stronger influence on employees’ psychological responses than technological strains. The paper offers a fresh perspective on employees’ psychological responses, digital transformation, and the impact of AI on the human factor in technology-rich work environments.
BackgroundConvolutional neural networks (CNNs) have emerged as powerful artificial intelligence tools for medical image analysis, demonstrating substantial improvements in disease detection, classification, and diagnostic support. Despite increasing evidence regarding their clinical performance, implementation within resource-constrained healthcare environments remains challenging due to limitations in infrastructure, computational resources, dataset availability, and healthcare workforce capacity. Understanding the current evidence relating to CNN architectures, performance characteristics, and deployment barriers is necessary to support sustainable implementation in underserved healthcare settings.AimThis review aimed to critically map the available evidence regarding convolutional neural networks for medical imaging in resource-constrained settings, with emphasis on CNN architectures, diagnostic performance, and deployment challenges.MethodsThis scoping review followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews (PRISMA-ScR) guideline. Literature searches were conducted across PubMed, Scopus, Web of Science, IEEE Xplore, and ScienceDirect from database inception to April 2026. Search terms combined Medical Subject Headings (MeSH) and free-text keywords using Boolean operators. Eligibility criteria were developed using the population–concept–context (PCC) framework. Peer-reviewed English-language studies investigating CNN applications in medical imaging relevant to resource-constrained healthcare settings were included. Data extraction and synthesis were performed thematically.ResultsA total of 18 studies met the inclusion criteria. Five major themes were identified: diagnostic accuracy and clinical performance of CNN models; lightweight and resource-efficient CNN architectures for low-resource deployment; explainability and interpretability of CNN systems; transfer learning and optimization for limited datasets; and implementation barriers affecting deployment in resource-constrained settings. CNN systems demonstrated strong diagnostic performance across tuberculosis, pneumonia, malaria, skin cancer, and other imaging applications, with several studies reporting near expert-level performance. Lightweight architectures and transfer learning approaches improved computational feasibility, while explainability methods enhanced transparency and clinician confidence. However, implementation challenges including computational limitations, poor connectivity, dataset scarcity, and algorithmic bias persisted across studies.ConclusionCNN-based medical imaging systems demonstrate considerable potential for strengthening diagnostic capacity within underserved healthcare environments. Nevertheless, sustainable implementation requires greater emphasis on context-specific model development, explainability, locally representative datasets, and real-world deployment studies to ensure equitable and effective adoption in resource-constrained healthcare systems.
Early disease detection is crucial for any disease it helps in better treatment and curing. For this purpose, several Artificial Intelligence (AI)-assisted research studies have been developed and tested for medical image processing. Although traditional Deep Learning (DL) models aid in disease detection, computational complexity and energy-intensive challenges limit the performance. For this reason, this research proposes SPiKe-Med, an efficient, accurate and neuromorphic-inspired solution by incorporating Spiking Neural Networks (SNNs) into state-of-the-art Convolutional Neural Network (CNN) and transformer backbones architectures. Stage 1 uses standardized preprocessing, N4 bias-field correction, isotropic resampling, z-score intensity normalization, and modality-specific augmentations. Stage 2 proposes Swin Unet-based Transformer-Network (SUT-Net), a high-performance dense segmentation backbone configured using benchmark Swin Transformer and UNET- Transformers (UNETR) for long-range context. Stage 3 develops the proposed SPiKe-Med by training multi-scale feature maps into spike trains with a learned encoder (temporal encoding) and feeds them to a spiking refinement network of Leaky Integrate-and-Fire (LIF) neurons trained as reconfigurable with SpikingJelly. Stage 4 trains the proposed SPiKe-Med on a two-track schedule: end-to-end surrogate-gradient optimization of SNN components (surrogate derivatives) and selective ANN to SNN porting of pretrained CNN encoder weights to ultra-low-latency inference. Loss functions combine a Dice, focal and temporal consistency term, which is optimized by temperature scaling. The proposed SPiKe-Med is tested with multiple benchmark datasets for different modalities like MRI and CT using Dice, Hausdorff distance, spike-rate and estimated energy per inference.
The complexity of data pertaining to patients diagnosed with cancer requires a shift from fragmented, unimodal diagnostics towards multimodal artificial intelligence (MAI) in order to achieve true precision oncology. In this literature review we examined the landscape of unimodal data modalities used in oncological practice including clinical records, multi-scale imaging (radiology and histopathology), and multi-omics signatures, alongside a critical comparison of the deep-learning architectures and integrative frameworks used to combine these data sources into advanced predictive models. By using fusion strategies (early, late, intermediate, and hybrid), MAI models are able to bridge the gap between genotype and phenotype, uncovering biological interactions that remain invisible to single-modality analysis. Current applications demonstrate significant improvements in diagnostic sensitivity, automated tumor grading, and the prediction of complex clinical outcomes, such as immunotherapy response and overall survival, referencing leading-edge tools and frameworks currently used or in active research. However, the transition from research to clinical practice is hindered by limitations such as data fragmentation, demographic biases, limited model explainability, and evolving regulatory requirements. We further outline emerging directions, including multimodal foundation models, large language models, and retrieval-augmented, agent-based systems. We concluded that the convergence of multimodal data streams and biologically informed AI represents the essential step for the next generation of personalized cancer care.
Wellbore trajectory deviation remains one of the major operational challenges encountered during directional and extended-reach drilling because even small departures from the planned well path can lead to poor reservoir placement, wellbore instability, increased non-productive time, and significant drilling costs. In most field operations, trajectory monitoring depends on periodic directional surveys together with threshold-based diagnostics. Although these methods are widely used, they often identify deviations only after they have become operationally noticeable, limiting the opportunity for timely corrective action. To overcome the limitations of conventional monitoring, the present study proposes an unsupervised deep learning framework capable of identifying the early onset of trajectory deviation by analyzing integrated well-log and geo-mechanical data without relying on labeled deviation events. The proposed framework combines Depth, Gamma Ray (GR), Shale Volume (Vsh), Resistivity, Sonic Transit Time (ΔT), P-wave Velocity (Vp), S-wave Velocity (Vs), Bulk Density, Calculated Density, Neutron Porosity (NPHI), Density Porosity (DPHI), and Poisson's Ratio to capture the lithological and mechanical characteristics that influence drilling behavior and trajectory stability. An LSTM Autoencoder (LSTM-AE) is employed to learn the normal temporal evolution of drilling parameters and identify anomalous behavior through reconstruction error. To complement the sequential learning capability of the autoencoder, a Graph Neural Network (GNN) is developed to represent the physical and geological relationships among the measured parameters, allowing complex multivariate interactions to be analyzed without requiring labeled datasets. The performance of both models is evaluated by comparing their ability to distinguish normal drilling behavior from progressively unstable operating conditions. The obtained results demonstrate that the LSTM-AE effectively learns the sequential characteristics of stable drilling and provides reliable early warning through changes in reconstruction error, whereas the GNN offers improved discrimination of trajectory-related anomalies by modeling the underlying relationships between geological and geo-mechanical variables. Collectively, these complementary approaches enableearlier recognition of developing trajectory deviations while reducing false alarms compared with conventional monitoring techniques. The findings demonstrate that integrating temporal sequence modeling with graph-based relational learning provides a practical and scalable solution for intelligent wellbore trajectory monitoring, supporting improved wellbore stability assessment, safer drilling operations, and more informed decision-making in complex subsurface environments.
Clinically useful artificial intelligence and machine-learning studies in healthcare require interpretable features, internal validation, and explicit boundaries between primary inference and external context. Among 1,612 baseline Parkinson’s Progression Markers Initiative (PPMI) participants, 1,439 had evaluable post-baseline cognition and contributed 5,909 records through Year 5. Composite cognitive progression occurred in 720 participants (50.0%). Each standard deviation increase in baseline mood/non-motor burden was associated with higher odds of progression (adjusted odds ratio 1.37, 95% confidence interval 1.20–1.55). The 1,000-resample participant bootstrap interval was 1.20–1.57, and estimates were stable across 1- to 5-year windows (odds ratios 1.35–1.41). In 10 repetitions of stratified 5-fold cross-validation, adding composite burden to clinical covariates produced a modest increase in mean area under the receiver operating characteristic curve from 0.665 to 0.682. We interpreted the PPMI result alongside separate neuroimaging, transcriptomic, and digital analyses in other samples and specified a future same-participant study; the separate analyses were not used for participant-level integration or validation. In the small resting-state functional magnetic resonance imaging cohort, most static and dynamic comparisons did not survive false-discovery-rate correction; two threshold-specific network-based-statistic components were retained as exploratory hypotheses. Molecular rankings were consistent with previously reported Parkinson’s disease biology, while wearable and voice datasets demonstrated feasibility for the source-task only. Baseline mood/non-motor assessment may support future risk-enrichment research, but the limited cross-validated increment, absence of external clinical validation, and lack of participant-matched multimodal data preclude clinical implementation.
The AI industry is advancing toward an Artificial General Intelligence (AGI) that will possess cognitive capabilities that are equal to or stronger than those of humans. While a purely computational AGI may be cognitively powerful, it will probably struggle to understand the human social environment. This gap needs both a theoretical and an applicative solution. This article provides the theoretical solution, specifying what such an AGI must contain and how it must be formed. We propose a theoretical construct of an AGI that will not only be computational but will also be capable to develop internal social understanding, internalize norms in a manner similar to humans, and act upon this social understanding. We call such AGI a “Social AGI.” We argue that Social AGI cannot be developed only by using technical alignment or external governance. Instead, it must undergo a staged development in which visible social behavior gradually becomes social cognition and norm internalization. This process cannot be separated from the parallel support that human society must provide the effort of developing Social AGI. While AI systems may become more familiar and embedded in institutions and society itself, the public perceptions of risk will probably decline. This normalization will provide the basis for granting the Social AGI progressively higher levels of autonomy. In order to capture this process of mutual change, we develop a socio-technical concept of human-AGI co-formation. We differentiate co-formation from co-existence, which describes humans and AI operating side by side, and from co-evolution, which describes reciprocal and cyclic influence. Co-formation refers to a process in which at the beginning Social AGI is shaped exclusively by human society. However, as the process progresses, they shape together the basic conditions under which the AGI is developed. On the AGI side, this will require the capabilities to internalize norms, and to be able to function according to its roles. On the societal side, this will require to create several mechanisms of role assignment, training and socializing the AGI, and monitoring its value alignment. The most important characteristics society should develop is a public acceptance and justification for increasing the AGI’s autonomy. Society may learn not only how to supervise the AGI, but also it may learn to trust it.
IntroductionKidney abnormalities, including cysts, tumors, and stones, are the most common renal disorders that can lead to severe complications such as chronic kidney disease or renal failure. Deep learning-based medical image analysis offers an effective approach for the accurate classification of kidney abnormalities, aiding the early diagnosis of renal disorders. However, its centralized training leads to inadequate privacy protection.MethodsConsidering the importance of ensuring individuals' data privacy, this study proposes a novel federated transfer learning framework for accurate classification of renal abnormalities using 12,446 kidney CT scan images and simultaneously preserves data privacy. CT scan images were preprocessed by resizing and normalization, followed by data augmentation techniques, including random rotations (±30°), horizontal flips, and color jitter, to address class imbalance and improve model generalization. Five pre-trained deep learning models such as MobileNetV2, EfficientNetV2-S, ResNet50, DenseNet121, and InceptionResNetV2 were trained across seven federated clients. Federated weighted averaging was employed for aggregation, and AES-256 encryption in CBC mode was applied to all model parameter transmissions between clients and the server.ResultsMobileNetV2 achieved the best performance, attaining 99.48% accuracy, 99.29% precision, 99.32% recall, 99.3% F1-score, 0.9999 AUC-ROC, and log loss of 0.0247. Cross-client validation produced an average accuracy of 98.85% with a generalization gap of only −0.0063, indicating strong generalization across client datasets.DiscussionThe proposed framework provides an effective balance between privacy preservation and communication efficiency, highlighting its potential for deployment in distributed clinical environments for kidney disease diagnosis.
BackgroundAccessing large-scale clinical and biomedical databases remains a significant barrier for clinicians and researchers, requiring substantial computational expertise. Agentic artificial intelligence frameworks, in which large language models (LLMs) orchestrate multi-step reasoning and query execution under interactive human supervision, offer the potential to democratize data access and accelerate evidence generation.MethodsWe applied the Model Context Protocol (MCP) to integrate, within a single agentic workflow, an Observational Medical Outcomes Partnership (OMOP)-standardized electronic health record (EHR) database from an academic health system (>7 million subjects), with the Scalable Precision Medicine Open Knowledge Engine (SPOKE), a curated knowledge graph integrating relationships biomedical concepts from expert-maintained resources. Using this implementation (MedCP), we evaluated the approach across 100 benchmarking clinical research tasks, 617 biomedical factual accuracy questions (BiomixQA), an integrative case study linking clinical co-occurrence with molecular similarity, and a standardized protocol for generating and replicating real-world studies.ResultsAcross the 100-task benchmark, knowledge-graph access improved mean scores overall for GPT-5.5 and Claude Opus 4.8, though the pattern differed between models; on BiomixQA, SPOKE grounding raised multiple-choice accuracy for both models, without changing true/false accuracy. In the case study, disease co-occurrence patterns extracted from the EHR correlated with molecular network similarity, surfacing mechanistic hypotheses from real-world data. Applied to the replication of published observational studies, the research protocol compressed timelines from months to hours, lowering the technical barrier to query generation and execution while study design and interpretation remained under expert supervision.ConclusionsAn agentic AI infrastructure that combines institutional EHR data with curated biomedical knowledge via MCP can serve as a transparent, domain-grounded, supervised research assistant for real-world evidence, supporting both hypothesis generation and testing.
Vision Transformers (ViTs) perform well on clean aerial imagery but degrade sharply when deployed in post-earthquake UAV operations, where motion blur, dust haze, illumination variation, and sensor noise combine to produce what we term post-earthquake visual drift, a structured distributional shift that can render an otherwise capable model dangerously unreliable in the field. Retraining or conventional domain adaptation is not a realistic option under the latency, compute, and annotation constraints of active disaster response. This work presents a comprehensive empirical evaluation of lightweight label-free test-time adaptation (TTA) methods for improving the robustness of ViTs under such deployment conditions, systematically comparing consistency-based self-supervision (MEMO) and entropy minimization (TENT) across datasets with varying classification complexity. Consistent with existing lightweight TTA approaches, only the LayerNorm affine parameters are adapted during inference, while the backbone and attention weights remain frozen. On the UAV-TEBDE post-earthquake dataset, accuracy rises from 73.47% under severe drift (severity 0.3) to 89.18% using MEMO, corresponding to a recovery of 15.71 percentage points over the drifted baseline, without using any ground-truth labels or performing any retraining. On UAV-TEBDE, MEMO also achieved higher recovery than the SAR baseline evaluated in this study. Additional evaluation on AID (30-class aerial scene classification) and UC Merced (21-class land-use classification) shows consistent drift-induced degradation and adaptation-driven recovery across datasets of varying complexity. The cross-dataset evaluation further reveals that the relative effectiveness of lightweight TTA methods depends on the classification-space dimensionality. TENT-based entropy minimization outperforms consistency-based MEMO when the class count is large, while MEMO is the stronger choice in low-class settings. This interaction between method design and classification-space dimensionality has direct consequences for choosing TTA strategies in operational disaster assessment pipelines. Feature-space PCA analysis confirms that LayerNorm-only updates geometrically recalibrate internal representations, rather than simply correcting the model's output logits.