Senior Observability Engineer with OpenTelemetry
Project overview
The project focuses on building a modern observability platform for AI driven applications and autonomous agent ecosystems. The solution provides comprehensive tracing, metrics collection, logging, and monitoring capabilities across cloud based services, enabling reliable operations, performance analysis, and operational insights.
Position overview
We are looking for a Senior Observability Engineer with strong experience in OpenTelemetry, AWS observability services, and Python based instrumentation. You will design and implement observability solutions for AI powered applications and agent based systems, ensuring end to end visibility, monitoring, tracing, and telemetry collection across distributed environments.
Team
You will work with observability engineers, software engineers, AI engineers, cloud specialists, and platform architects. The team collaborates closely on designing scalable telemetry solutions, defining observability standards, and supporting production grade AI and cloud workloads.
Technology stack
Python, OpenTelemetry, AWS Distro for OpenTelemetry (ADOT), AWS CloudWatch, AWS X Ray, AWS AgentCore Observability, LangChain, W3C Trace Context, W3C Baggage Propagation, YAML, Git, CI/CD, Cloud Infrastructure
Responsibilities
- Develop and maintain observability solutions for cloud native and AI powered applications
- Implement OpenTelemetry instrumentation for Python services and distributed systems
- Design telemetry pipelines using AWS Distro for OpenTelemetry
- Configure and optimize metrics, logs, and tracing integrations with AWS CloudWatch and AWS X Ray
- Implement distributed tracing using W3C Trace Context and Baggage propagation standards
- Integrate and monitor LangChain based workflows and AI application components
- Design structured logging standards and telemetry data models
- Collaborate with engineering teams to define observability requirements and establish monitoring best practices
- Support performance analysis, troubleshooting, and incident investigation activities
- Contribute to the evolution of observability architecture for agent based systems and AI workloads
Requirements
- 3+ years of experience in observability engineering
- Experience implementing OpenTelemetry instrumentation for Python applications and services
- Experience with AWS CloudWatch including Metrics, Logs, and Structured Logs
- Experience integrating AWS X Ray for distributed tracing and performance analysis
- Strong understanding of telemetry concepts including metrics, traces, logs, and distributed observability
- Experience working with cloud environments and production systems
- Knowledge of W3C Trace Context and Baggage propagation standards
- Experience working with source control systems and collaborative development practices
Nice to have
- Experience with LangGraph or Strands agent frameworks
- Knowledge of OpenTelemetry GenAI SIG semantic conventions
- Experience configuring AWS Distro for OpenTelemetry collectors and export pipelines
- Familiarity with AWS AgentCore Observability capabilities
- Experience supporting observability for AI applications, agentic systems, or LLM powered solutions
- Knowledge of cloud native monitoring and distributed systems architectures