All rolesDevOps & Cloud
Site Reliability Engineer (SRE) A deep-dive guide for SREs and Platform Engineers, focusing on Error Budgets, SLIs/SLOs, Observability, Kubernetes orchestration, Incident Response, and system scalability.
336 questions Updated 2026-02-08 Beginner Intermediate Advanced
What you will be asked about SRE Fundamentals & Principles Monitoring & Observability Alerting & On-Call Incident Management System Design & Architecture Capacity Planning & Performance Kubernetes & Container Orchestration Linux & Operating Systems Automation & Infrastructure as Code Databases & Data Systems Cloud Platforms & Distributed Systems Security & Compliance Scenario-Based Questions Behavioral & Leadership Real Production Scenarios
How to prepare Go through the topic list above and mark every one you cannot explain for five minutes unprepared. Those are your gaps. Pair every concept with a story from your own work — interviewers probe depth, and depth comes from having actually done it. Do the DSA rounds anyway. Almost every role in this list still screens with coding. Prepare two projects you can whiteboard end to end, including what you would change now. Site Reliability Engineer (SRE) interview questions336 1. What is Site Reliability Engineering? Beginner 2. What is the difference between SRE and DevOps? Intermediate 3. What is the difference between SRE and traditional operations? Intermediate 4. What are the main responsibilities of an SRE? Beginner 5. What is the error budget? Intermediate 6. How do you calculate error budget? Intermediate 7. What happens when error budget is exhausted? Intermediate 8. What is SLI (Service Level Indicator)? Beginner 9. What is SLO (Service Level Objective)? Beginner 10. What is SLA (Service Level Agreement)? Beginner 11. What is the difference between SLI, SLO, and SLA? Intermediate 12. Give examples of good SLIs for a web service. Beginner 13. How do you set SLOs? Intermediate 14. What is the recommended SLO target (e.g., 99.9% vs 99.99%)? Intermediate 15. What is toil in SRE? Beginner 16. How do you measure toil? Intermediate 17. What is the acceptable percentage of toil for SREs? Beginner 18. How do you reduce toil? Intermediate 19. What is the 50% engineering time rule in SRE? Intermediate 20. What is blameless postmortem? Beginner 21. What should be included in a postmortem document? Intermediate 22. What is the difference between incident and outage? Beginner 23. What is mean time to detect (MTTD)? Beginner 24. What is mean time to repair/recover (MTTR)? Beginner 25. What is mean time between failures (MTBF)? Intermediate 26. What is the difference between monitoring and observability? Intermediate 27. What are the three pillars of observability? Beginner 28. What is the difference between metrics, logs, and traces? Intermediate 29. What is the four golden signals of monitoring? Intermediate 30. Explain latency as a golden signal. Beginner 31. Explain traffic as a golden signal. Beginner 32. Explain errors as a golden signal. Beginner 33. Explain saturation as a golden signal. Intermediate 34. What metrics would you monitor for a web application? Beginner 35. What metrics would you monitor for a database? Intermediate 36. What metrics would you monitor for a Kubernetes cluster? Intermediate 37. What is white-box monitoring vs black-box monitoring? Intermediate 38. What is Prometheus and how does it work? Beginner 39. What is the Prometheus data model? Intermediate 40. What is PromQL? Intermediate 41. Write a PromQL query to calculate request rate. Intermediate 42. Write a PromQL query to calculate error rate. Advanced 43. Write a PromQL query for p95 latency. Advanced 44. What is Grafana and how do you use it with Prometheus? Beginner 45. What is the pushgateway in Prometheus? Intermediate 46. What is service discovery in Prometheus? Intermediate 47. What is the ELK/EFK stack? Beginner 48. What is structured logging vs unstructured logging? Intermediate 49. What is distributed tracing? Intermediate 50. What is Jaeger or Zipkin? Beginner 51. What is span and trace in distributed tracing? Intermediate 52. What is sampling in tracing? Advanced 53. How do you correlate logs, metrics, and traces? Advanced 54. What is cardinality in monitoring and why does it matter? Advanced 55. What is alert fatigue and how do you prevent it? Beginner 56. What makes a good alert? Intermediate 57. What is the difference between alert and notification? Beginner 58. What are the alert severity levels you use? Beginner 59. How do you write actionable alerts? Intermediate 60. What is alert routing? Beginner 61. What is escalation policy? Intermediate 62. What is PagerDuty/OpsGenie/VictorOps? Beginner 63. How do you handle alert fatigue? Intermediate 64. What is alert suppression vs alert inhibition? Advanced 65. What is flapping in alerting? Intermediate 66. How do you reduce false positives? Intermediate 67. What is on-call rotation? Beginner 68. What is a reasonable on-call schedule? Intermediate 69. How do you handle burnout from on-call? Intermediate 70. What do you do when you get paged at 3 AM? Beginner 71. What is runbook automation? Intermediate 72. What should be in an on-call runbook? Beginner 73. Difference between page-worthy and non-page-worthy alerts? Intermediate 74. How do you prioritize multiple simultaneous incidents? Advanced 75. What is alert aggregation? Intermediate 76. What is maintenance window and how do you handle alerts during it? Beginner 77. How do you measure on-call quality? Intermediate 78. What is time to acknowledge (TTA)? Beginner 79. What is time to resolve (TTR)? Beginner 80. How do you improve MTTR? Intermediate 81. What is an incident? Beginner 82. What are incident severity levels? Beginner 83. What is SEV-1, SEV-2, SEV-3 incidents? Beginner 84. What is the incident response process? Intermediate 85. Roles in incident response (IC, scribe, communications)? Intermediate 86. What is an Incident Commander (IC)? Intermediate 87. What are the responsibilities of IC? Intermediate 88. What is incident communication strategy? Advanced 89. How do you communicate during an outage? Beginner 90. What is a status page? Beginner 91. What is incident timeline? Intermediate 92. How do you conduct an incident call/war room? Intermediate 93. What is incident escalation? Beginner 94. When do you escalate an incident? Intermediate 95. What is incident mitigation vs resolution? Intermediate 96. What is rollback vs roll-forward during incident? Intermediate 97. What is incident retrospective/postmortem? Beginner 98. What is blameless culture? Beginner 99. What are the 5 whys technique? Intermediate 100. What is root cause analysis (RCA)? Beginner 101. What is contributing factor vs root cause? Advanced 102. What should be in a postmortem document? Beginner 103. What are action items from postmortem? Intermediate 104. How do you track action items? Intermediate 105. How do you prevent incident recurrence? Intermediate 106. What is incident review meeting? Beginner 107. How do you learn from incidents? Beginner 108. Chaos engineering and how it relates to incidents? Advanced 109. What is game day exercise? Intermediate 110. What is disaster recovery drill? Advanced 111. How do you design a highly available system? Intermediate 112. What is the difference between high availability and fault tolerance? Intermediate 113. What is redundancy? Beginner 114. What is active-active vs active-passive architecture? Intermediate 115. What is load balancing? Beginner 116. What are load balancing algorithms? Intermediate 117. What is health check in load balancing? Beginner 118. What is circuit breaker pattern? Advanced 119. What is retry logic and exponential backoff? Intermediate 120. What is rate limiting? Beginner 121. What is throttling vs rate limiting? Intermediate 122. What is caching strategy? Intermediate 123. What is cache invalidation? Intermediate 124. What is CDN and when to use it? Beginner 125. What is database replication? Intermediate 126. Master-slave vs multi-master replication? Intermediate 127. What is database sharding? Advanced 128. What is horizontal vs vertical scaling? Beginner 129. What is stateless vs stateful application? Intermediate 130. How do you design for failure? Intermediate 131. What is graceful degradation? Intermediate 132. What is bulkhead pattern? Advanced 133. What is timeout strategy? Intermediate 134. What is idempotency and why is it important? Advanced 135. What is eventual consistency? Advanced 136. What is CAP theorem? Advanced 137. How do you handle single point of failure (SPOF)? Beginner 138. What is disaster recovery (DR)? Beginner 139. What is RTO (Recovery Time Objective)? Intermediate 140. What is RPO (Recovery Point Objective)? Intermediate 141. What is capacity planning? Beginner 142. How do you forecast capacity needs? Intermediate 143. What is the difference between load and capacity? Beginner 144. What is headroom in capacity planning? Intermediate 145. What is utilization vs saturation? Advanced 146. How do you measure system capacity? Intermediate 147. What is performance testing? Beginner 148. What is load testing vs stress testing? Intermediate 149. What is spike testing? Intermediate 150. What is soak testing (endurance testing)? Intermediate 151. Tools for load testing (JMeter, Gatling, Locust)? Beginner 152. What is latency budget? Advanced 153. What is throughput? Beginner 154. What is the difference between latency and throughput? Beginner 155. What is queueing theory in SRE? Advanced 156. What is Little's Law? Advanced 157. How do you optimize database performance? Intermediate 158. How do you optimize API performance? Intermediate 159. What is connection pooling? Intermediate 160. What is database query optimization? Intermediate 161. What is indexing strategy? Intermediate 162. What is N+1 query problem? Advanced 163. How do you identify performance bottlenecks? Intermediate 164. What is profiling? Advanced 165. What is the USE method (Utilization, Saturation, Errors)? Advanced 166. What is Kubernetes in SRE context? Beginner 167. What Kubernetes metrics do you monitor? Intermediate 168. What is pod crash loop backoff? Beginner 169. What is OOMKilled error? Intermediate 170. How do you troubleshoot pending pods? Intermediate 171. What is resource requests vs limits? Intermediate 172. What is HPA (Horizontal Pod Autoscaler)? Beginner 173. What is VPA (Vertical Pod Autoscaler)? Advanced 174. What is cluster autoscaler? Intermediate 175. What is liveness probe vs readiness probe? Intermediate 176. How do you set probe thresholds? Intermediate 177. What is PodDisruptionBudget (PDB)? Advanced 178. What is node affinity and pod affinity? Advanced 179. What is taints and tolerations? Advanced 180. How do you perform rolling updates safely? Intermediate 181. What is deployment strategy (RollingUpdate, Recreate)? Beginner 182. How do you rollback a deployment? Beginner 183. What is StatefulSet and when to use it? Intermediate 184. What is DaemonSet use case? Beginner 185. What is resource quota and limit range? Advanced 186. How do you monitor Kubernetes cluster health? Intermediate 187. What is kube-state-metrics? Intermediate 188. What is node exporter? Beginner 189. What is kubectl top command? Beginner 190. How do you debug a container? Beginner 191. What is ephemeral containers? Advanced 192. What is Kubernetes Events? Intermediate 193. How do you handle persistent storage in K8s? Intermediate 194. What is StorageClass? Intermediate 195. What is CNI (Container Network Interface)? Advanced 196. How do you check system load average? Beginner 197. What does load average 1, 5, 15 mean? Intermediate 198. How do you identify high CPU usage process? Beginner 199. How do you identify high memory usage? Beginner 200. What is the difference between memory and swap? Beginner 201. What is OOM killer? Intermediate 202. How do you troubleshoot disk space issues? Intermediate 203. What is inode and how can you run out of inodes? Advanced 204. How do you find which process is using a file? Beginner 205. What is lsof command? Intermediate 206. How do you check network connections? Beginner 207. What is netstat/ss command? Beginner 208. How do you troubleshoot DNS issues? Intermediate 209. Trace network packets (tcpdump, wireshark)? Intermediate 210. What is strace and when do you use it? Advanced 211. What is system calls? Intermediate 212. How do you check process threads? Intermediate 213. What is context switching? Advanced 214. What is soft vs hard limits (ulimit)? Intermediate 215. How do you tune kernel parameters (sysctl)? Advanced 216. What is a file descriptor and how to increase limits? Intermediate 217. What is TCP time_wait state? Advanced 218. What is TCP connection states? Advanced 219. How do you troubleshoot performance using perf? Advanced 220. What is eBPF and its use in SRE? Advanced 221. How do you automate toil? Beginner 222. What is configuration management (Ansible, Puppet, Chef)? Beginner 223. What is infrastructure as code (Terraform, CloudFormation)? Beginner 224. What is the difference between imperative and declarative IaC? Intermediate 225. What is GitOps? Intermediate 226. What is reconciliation loop? Advanced 227. How do you manage secrets in automation? Intermediate 228. What is idempotency in automation? Intermediate 229. How do you test infrastructure code? Advanced 230. What is policy as code? Advanced 231. What is continuous deployment vs continuous delivery? Intermediate 232. How do you implement safe deployment practices? Intermediate 233. What is canary deployment? Beginner 234. What is blue-green deployment? Intermediate 235. What is feature flag? Beginner 236. How do you automate rollback? Advanced 237. What is progressive delivery? Advanced 238. What is deployment pipeline? Beginner 239. How do you automate incident response? Advanced 240. What is ChatOps? Beginner 241. How do you monitor database health? Beginner 242. What is connection pool exhaustion? Intermediate 243. How do you troubleshoot slow queries? Intermediate 244. What is query execution plan? Advanced 245. How do you handle database failover? Intermediate 246. What is replication lag? Beginner 247. What is split-brain problem in databases? Advanced 248. How do you perform database backup and restore? Intermediate 249. What is point-in-time recovery (PITR)? Advanced 250. How do you test database backups? Intermediate 251. What cloud platforms have you worked with? Beginner 252. How do you ensure reliability in cloud? Intermediate 253. What is multi-AZ deployment? Beginner 254. What is multi-region deployment? Intermediate 255. How do you handle cloud provider outages? Advanced 256. What is AWS Auto Scaling? Beginner 257. What is ELB health checks? Beginner 258. How do you monitor cloud resources? Intermediate 259. What is CloudWatch vs Prometheus for cloud monitoring? Intermediate 260. What is distributed system? Beginner 261. What are challenges in distributed systems? Intermediate 262. What is network partition? Intermediate 263. What is split-brain scenario? Advanced 264. How do you handle eventual consistency? Advanced 265. What is distributed consensus (Raft, Paxos)? Advanced 266. What is service mesh (Istio, Linkerd)? Intermediate 267. What is sidecar pattern? Intermediate 268. What is API gateway? Beginner 269. What is backpressure in distributed systems? Advanced 270. How do you handle cascading failures? Advanced 271. What is bulkhead pattern in microservices? Advanced 272. What is timeout propagation? Advanced 273. What is distributed tracing importance? Intermediate 274. How do you debug distributed systems? Intermediate 275. What is observability in microservices? Intermediate 276. What is security in SRE role? Beginner 277. How do you implement least privilege principle? Intermediate 278. What is secrets management? Intermediate 279. What is certificate rotation? Intermediate 280. How do you handle security incidents? Intermediate 281. What is DDoS mitigation? Intermediate 282. What is rate limiting for security? Beginner 283. How do you monitor for security threats? Intermediate 284. What is audit logging? Beginner 285. What is compliance monitoring? Intermediate 286. How do you handle vulnerability patching? Intermediate 287. What is patch management strategy? Intermediate 288. What is security scanning in CI/CD? Intermediate 289. How do you ensure data encryption? Beginner 290. What is principle of defense in depth? Intermediate 291. Website is slow - how do you troubleshoot? Intermediate 292. Database is down - what are your steps? Intermediate 293. CPU is at 100% - how do you investigate? Intermediate 294. Memory is exhausted - what do you do? Intermediate 295. Disk is full - how do you handle it? Beginner 296. Pods are crash looping - how do you debug? Beginner 297. Load balancer shows unhealthy targets - what do you check? Intermediate 298. Latency suddenly increased - how do you investigate? Intermediate 299. Error rate spiked - what are your steps? Beginner 300. Traffic dropped to zero - how do you troubleshoot? Intermediate 301. Deployment failed - how do you rollback? Beginner 302. Database replication lag is high - what do you do? Intermediate 303. You're paged for high memory usage - what do you do? Intermediate 304. SSL certificate expired - how do you handle? Beginner 305. DNS resolution failing - how do you troubleshoot? Intermediate 306. Application throwing 500 errors - how do you debug? Beginner 307. Kafka consumer lag increasing - what do you check? Advanced 308. Redis cache hit rate dropped - what do you investigate? Intermediate 309. Network latency between services increased - how do you debug? Advanced 310. Cloud cost suddenly increased - how do you investigate? Intermediate 311. How would you design monitoring for a new service? Intermediate 312. How would you implement zero-downtime deployment? Intermediate 313. How would you handle a complete datacenter failure? Advanced 314. How would you migrate a service with zero downtime? Advanced 315. How would you handle Black Friday traffic spike? Intermediate 316. How would you implement disaster recovery? Intermediate 317. Reduce deployment time from 30 mins to 5 mins? Advanced 318. How would you improve MTTR for your team? Intermediate 319. Handle runaway process consuming resources? Intermediate 320. Debug intermittent timeout issues? Advanced 321. Service is healthy but receiving no traffic - what do you check? Intermediate 322. Memory leak - identify cause? Intermediate 323. Database connections maxed out - what are your steps? Intermediate 324. How do you handle third-party API outage? Intermediate 325. How would you design SLOs for a payment service? Intermediate 326. How would you reduce toil in your current role? Beginner 327. Handle competing incidents simultaneously? Intermediate 328. Onboard a new service to your monitoring stack? Beginner 329. How would you implement chaos engineering? Advanced 330. Improve observability of legacy system? Advanced 331. Tell me about your most challenging production incident. Intermediate 332. How do you prioritize during an outage? Beginner 333. Describe a time you prevented a major incident. Intermediate 334. How do you handle stress during critical incidents? Beginner 335. How do you balance reliability and feature velocity? Intermediate 336. Post-deploy performance regression - identify cause? Advanced