Sunnyvale, CA, USA
2020
  |  By Sejal Pandey
Set up the NVIDIA DCGM Exporter with Docker or Helm, pick the right DCGM metrics, enable profiling counters, map GPUs to Kubernetes pods, and add alerts. Sejal Pandey works on content and growth at Last9, writing about observability, reliability, and SRE practices.
  |  By Sejal Pandey
A practical comparison of the best observability tools in 2026: Last9, Datadog, Grafana Cloud, New Relic, Dynatrace, Honeycomb, Coralogix and SigNoz. Sejal Pandey works on content and growth at Last9, writing about observability, reliability, and SRE practices.
  |  By Sejal Pandey
An error budget is the unreliability an SLO allows. Learn to calculate it, track burn rate, set multiwindow alerts in PromQL and write an error budget policy. Sejal Pandey works on content and growth at Last9, writing about observability, reliability, and SRE practices.
  |  By Sejal Pandey
Grafana Cloud pricing broken down: free tier limits, Pro rates for metrics, logs and traces, how active series and DPM are billed, and worked cost examples. Sejal Pandey works on content and growth at Last9, writing about observability, reliability, and SRE practices.
  |  By Sejal Pandey
How to monitor a Rails application in production: reading Puma's stats endpoint, spotting GC pressure, and finding the query that's actually slow. Sejal Pandey works on content and growth at Last9, writing about observability, reliability, and SRE practices.
  |  By Sejal Pandey
SRE automation covers four different jobs: runbook automation, self-healing infrastructure, drift detection, and resilience testing. Sejal Pandey works on content and growth at Last9, writing about observability, reliability, and SRE practices.
  |  By Sejal Pandey
AI agent observability means tracing tool calls, reasoning steps, and handoffs, not just tokens and latency. Here's what to instrument and why. Sejal Pandey works on content and growth at Last9, writing about observability, reliability, and SRE practices.
  |  By Sejal Pandey
How to monitor a PHP application in production: reading the PHP-FPM status page, checking OPcache health, and finding the request that is actually slow. Sejal Pandey works on content and growth at Last9, writing about observability, reliability, and SRE practices.
  |  By Sejal Pandey
AIOps explained in plain terms: what it means, how it differs from AI-SRE and MLOps, how it detects problems, and whether a small team needs it. Sejal Pandey works on content and growth at Last9, writing about observability, reliability, and SRE practices.
  |  By Sejal Pandey
Eight LLM and AI agent observability platforms compared: what each tracks, pricing and free tiers, self-hosting options, and who each is built for. Sejal Pandey works on content and growth at Last9, writing about observability, reliability, and SRE practices.
  |  By Last9 - Monitoring for AI Native SDLC
View Kubernetes events in Last9 — across clusters, deployments, statefulsets, and even correlated with services.
  |  By Last9 - Monitoring for AI Native SDLC
How are you teaching your agents? Are they learning on their own? Does that lead to better results? Listen to @prathameshsonpatki7217 talk about our experience running agents in production for the last 8 months that self improve!
  |  By Last9 - Monitoring for AI Native SDLC
Platform engineering provides powerful tools that handle a lot under the hood. Learn how to calculate your remaining error budget with a simple formula using real numbers and objective statements.
  |  By Last9 - Monitoring for AI Native SDLC
30 minutes of eating crow! Learn from our SLO mistakes at Weave. Discover pitfalls and shortcuts to doing it right the first time. Avoid our wrong, wrong, wrong, wrongs!
  |  By Last9 - Monitoring for AI Native SDLC
OpenTelemetry aims to link metrics to traces and logs, offering OpenCensus users a seamless migration path. Work with existing protocols like Prometheus. Leverage existing tooling without learning something completely new.
  |  By Last9 - Monitoring for AI Native SDLC
OpenTelemetry explained: standards, SDKs for various languages (Ruby, Python, Go), and middleware tools. Deploy these to pre-process data and send it to your destination.
  |  By Last9 - Monitoring for AI Native SDLC
Stop debugging infrastructure issues across multiple dashboards. See how Last9's Discover Infrastructure monitors K8s pods and traditional hosts together—with resource analysis, pod-level debugging, and AI that correlates app problems to infrastructure root causes. One setup (K8s + host monitoring) → Complete infrastructure visibility that connects to your services and jobs. No more blind spots between application performance and underlying resources.
  |  By Last9 - Monitoring for AI Native SDLC
Stop debugging background jobs with docker logs and prayer. See how Last9's Discover Jobs monitors async operations like APIs—with P95 latencies, error breakdowns, and operation-level traces for every job type.

Last9 provides tools to improve Reliability in large-scale cloud-native environments.

Our open-standards-based tools provide visibility into the Rube Goldberg of micro-services. We take away the toil of managing a time series database by dramatically reducing your costs and improving developer productivity.

Levitate is our time series metrics & events warehouse designed for scale and high cardinality. Our warehousing capabilities provide necessary control levers to ensure cost-efficient data growth management, surpassing traditional storage solutions.

Start your observability journey today with Levitate. A Managed Time Series Data Warehouse that SREs trust.