The stochastic nature of Large Language Models (LLMs) challenges traditional evaluation paradigms, which rely on single-response metrics and often mask complex behavioral patterns. This paper introduces Trait and Consistency Evaluation for LLMs (TraCE-LLM), an evaluation protocol that quantifies latent behavioral traits and model consistency within a black-box paradigm. Through a factorial design combining five LLMs, three benchmarks and a systematic stratification by prompt style (Naive, Chain-of-Thought and Adversarial), the framework employs a multidimensional rubric to measure Depth of Reasoning (DoR) and Originality (ORI) of model responses. The primary empirical contribution of this study is the identification and formalization of the Adversarial Compensation Effect (ACE), a phenomenon wherein smaller-capacity models under adversarial stress exhibit a paradoxical gain in accuracy metrics while suffering a severe degradation in behavioral stability. Our results also demonstrate an asymmetric stability with DoR being a significantly more stable trait than ORI and the prevalence of compressed reasoning, where 17.8% of correct answers lack adequate justification. By decoupling response correctness from process quality, TraCE-LLM provides a blueprint for more granular and reliable evaluation, arguing that LLM auditing must be multidimensional, context-sensitive and psychometrically informed to ensure the development of safer and more interpretable AI.