Degree

Doctor of Philosophy (PhD)

Department

Division of Computer Science and Engineering

Document Type

Dissertation

Abstract

Cloud microservices architectures have become the dominant paradigm for building large-scale applications due to their scalability, modularity, and deployment flexibility. However, maintaining predictable response time remains a fundamental challenge. Modern microservices frequently experience severe long latency spikes, even under moderate resource utilization, because transient performance disruptions propagate through complex execution dependencies. Understanding the underlying causes of response time instability and developing effective techniques for its diagnosis and mitigation are essential for building reliable cloud systems.   In this dissertation, we present a comprehensive study of response time instability in cloud microservices through the unified perspective of execution dependencies and resource contention. First, we characterize how transient bottlenecks propagate and amplify into significant end-to-end latency through blocking effects and transport-layer retransmissions. These findings demonstrate that response time instability is fundamentally a dependency propagation problem rather than merely a consequence of sustained hardware resource saturation. Second, we present SyncM and Grunt, two dependency-aware attacks that exploit execution dependencies and blocking effects to amplify transient performance disruptions into severe end-to-end latency degradation using low-volume request streams, revealing a previously unexplored class of performance attacks against cloud microservices. Third, we introduce PathFence, a path-aware resource isolation framework that partitions soft resources according to execution paths to mitigate cross-path blocking and significantly improve response time stability under resource contention (e.g., under bursty workloads and adversarial conditions). Finally, we present SORAT, a fine-grained soft-resource-aware tracing framework that explicitly captures worker and connection acquisition delays, constructs Soft-Resource-Aware Trace Graphs, and enables more accurate anomaly detection and root-cause localization by exposing hidden soft-resource contention and its propagation.   Collectively, these contributions provide a unified framework for understanding the causes, exploits, mitigations, and diagnosis of response time instability in cloud microservices, advancing the design of more reliable, resilient, and observable cloud applications.

Date

7-17-2026

Committee Chair

Wang, Qingyang

LSU Acknowledgement

1

LSU Accessibility Acknowledgment

1

Available for download on Sunday, July 15, 2029

Share

COinS