Degree
Doctor of Philosophy (PhD)
Department
Division of Computer Science and Engineering
Document Type
Dissertation
Abstract
Cloud microservices architectures have become the dominant paradigm for building large-scale applications due to their scalability, modularity, and deployment flexibility. However, maintaining predictable response time remains a fundamental challenge. Modern microservices frequently experience severe long latency spikes, even under moderate resource utilization, because transient performance disruptions propagate through complex execution dependencies. Understanding the underlying causes of response time instability and developing effective techniques for its diagnosis and mitigation are essential for building reliable cloud systems. In this dissertation, we present a comprehensive study of response time instability in cloud microservices through the unified perspective of execution dependencies and resource contention. First, we characterize how transient bottlenecks propagate and amplify into significant end-to-end latency through blocking effects and transport-layer retransmissions. These findings demonstrate that response time instability is fundamentally a dependency propagation problem rather than merely a consequence of sustained hardware resource saturation. Second, we present SyncM and Grunt, two dependency-aware attacks that exploit execution dependencies and blocking effects to amplify transient performance disruptions into severe end-to-end latency degradation using low-volume request streams, revealing a previously unexplored class of performance attacks against cloud microservices. Third, we introduce PathFence, a path-aware resource isolation framework that partitions soft resources according to execution paths to mitigate cross-path blocking and significantly improve response time stability under resource contention (e.g., under bursty workloads and adversarial conditions). Finally, we present SORAT, a fine-grained soft-resource-aware tracing framework that explicitly captures worker and connection acquisition delays, constructs Soft-Resource-Aware Trace Graphs, and enables more accurate anomaly detection and root-cause localization by exposing hidden soft-resource contention and its propagation. Collectively, these contributions provide a unified framework for understanding the causes, exploits, mitigations, and diagnosis of response time instability in cloud microservices, advancing the design of more reliable, resilient, and observable cloud applications.
Date
7-17-2026
Recommended Citation
Gu, Xuhang, "Response Time Instability in Cloud Microservices: Causes, Exploits, and Mitigations" (2026). LSU Doctoral Dissertations. 7141.
https://repository.lsu.edu/gradschool_dissertations/7141
Committee Chair
Wang, Qingyang
LSU Acknowledgement
1
LSU Accessibility Acknowledgment
1