Skip to main content
PrismCV
JobsExtensionPricing
LoginCheck Your Resume
Check Your Resume
← Back to all jobs

Principal Software Engineer - Observability & Telemetry Data

Smartsheet · Bellevue, WA, USA

On-site
Lead
Engineering - Developers

Posted September 1, 2026

Check your resume against this jobApply on Greenhouse

Job Description

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.

Observability Engineering owns the breadth and depth of observability and telemetry at Smartsheet. We build and operate the shared platform every engineering team depends on to understand how their software behaves in production (metrics, logs, distributed traces, RUM, synthetics, and AI/agentic telemetry) across a global, multi-region, commercial and government footprint. We are consolidating ownership of all of it into a single correlated platform, and treating telemetry as what it actually is: one of the largest and most valuable datasets at the company. It's a broad mandate with real scope to shape, and we hold this internal platform to the same rigor, reliability, and product standards as customer-facing software. 

As a Principal Software Engineer (Observability & Telemetry Data), you will set the technical direction for how Smartsheet collects, models, stores, and derives value from telemetry data. This is an architect role: you will define the telemetry data platform and the standards that govern it, carry that architecture across engineering org boundaries, partner with our Data and AI Platform teams on Databricks and MLflow integration, and act as the company's definitive technical voice on observability. You will influence roadmaps you do not own, resolve the hard design trade-offs between fidelity and cost, and raise the technical bar for every engineer instrumenting a service. 

This full-time position reports to the Team Lead, Observability Engineering and can be located in our Bellevue, WA office, or you may work remotely from anywhere in the US where Smartsheet is a registered employer. 

You Will 

  • Architect the Telemetry Data Platform: Own the end-to-end design for how telemetry is landed, modeled, and queried, including open table formats, partitioning and schema evolution strategy, separation of storage and compute, tiered retention, and the query interfaces engineers actually use. Make telemetry a durable, portable, Smartsheet-owned dataset rather than a vendor-locked byproduct. 
  • Own the Telemetry Data Model: Define the semantic conventions, shared business identifiers (user, org, plan, tenant), and schema standards that let any signal be correlated with any other, and drive their adoption across every service team at Smartsheet. 
  • Set OpenTelemetry Direction: Lead the migration to OTel-based instrumentation, defining collector architecture, context propagation, and sampling strategy (including tail-based sampling) so that engineers can move from a log line to a trace to a metric without losing the thread.
  • Integrate AI and Agentic Telemetry with Databricks: Own the architecture connecting our observability platform to Databricks and MLflow, so that agentic and model telemetry (prompt, completion, tool and MCP calls, evaluation results) is captured un-sampled, stays useful under the input/output redaction our governance requires, and reconciles cleanly with the traces and metrics in our primary observability stack. 
  • Instrument the Data Platform Itself: Bring first-class observability to our data estate, including Databricks jobs, pipelines, and warehouses, with meaningful signals for freshness, data quality, lineage, and cost, so data reliability is measured with the same discipline as service reliability. 
  • Build the Analytics Layer on Telemetry: Turn telemetry into decision-grade analytics, covering reliability and incident metrics, telemetry cost and chargeback models, and adoption and coverage reporting that leadership can act on. 
  • Engineer Collection and Routing at Scale: Architect the high-volume collection and routing tier (FluentBit, Kinesis, and OTel collectors) that moves telemetry from every service to its destination across US, EU, AU, and GovCloud regions, and own the migration of these pipelines as we consolidate onto a unified backend. 
  • Own Telemetry Economics: Set the cost architecture for observability data, including ingest governance, cardinality control, and storage tiering, so that teams get the fidelity they need to debug without the spend pressure that causes them to under-instrument. 
  • Carry Architecture Across Org Boundaries: Partner with the Data Platform, AI Platform, and infrastructure organizations to align telemetry architecture with theirs, influence roadmaps you do not own, and represent observability in company-level platform and vendor decisions. 
  • Raise the Technical Bar: Lead design and code reviews, author the architecture decisions and standards others build against, and mentor senior and mid-level engineers on instrumentation, telemetry data modeling, and cost-aware design. 
  • Participate in a production support and on-call rotation, taking ownership of the most complex issue resolution and driving root-cause analysis that improves system resiliency. 

You Have 

  • 10+ years of experience building and operating large-scale distributed systems, data platforms, or observability infrastructure, including time at Principal or Staff level. 
  • Deep Observability Expertise: Hands-on production ownership of metrics, logs, and distributed tracing at scale, including at least one major backend (Datadog, or comparable) and a working understanding of cardinality and cost mechanics. 
  • Telemetry Data Engineering: Demonstrated depth in large-scale data architecture, including open table formats (Delta Lake, Apache Iceberg), Spark or comparable distributed processing, streaming ingestion, partitioning and schema evolution, and query performance and cost tuning over very large datasets. 
  • Databricks Depth: Practical experience with the Databricks platform (jobs, clusters, Unity Catalog, Delta) and with MLflow for model and agent telemetry. 
  • OpenTelemetry Depth: Practical experience with OTel collectors, semantic conventions, context propagation, and sampling strategy, including tail-based sampling. 
  • Pipeline Engineering: Experience with high-volume log and telemetry pipelines, including FluentBit or Fluentd, streaming transport such as Kinesis or Kafka, and search backend index and mapping design.
  • Advanced AWS & Kubernetes Expertise: EKS, ECS Fargate, EC2, Lambda, and CloudWatch in production. 
  • 10+ years of programming experience with modern languages such as Go, Java, Python, or Scala, and strong SQL. 
  • Infrastructure as Code: Terraform, and GitOps workflows such as Flux or ArgoCD. 
  • Architectural Influence at Scale: A track record of setting technical direction that multiple teams and organizations adopted, including the written artifacts (architecture decisions, standards, RFCs) that made it durable after you moved on. 
  • Strong incident response instincts, with experience improving mean-time-to-resolution through better instrumentation and better data rather than more heroics. 
  • A degree in Computer Science, Engineering, or a related field, or equivalent practical industry experience.
  • Legally eligible to work in the U.S. on an ongoing basis 

Nice to Have 

  • Experience instrumenting LLM or agentic systems, including OTel GenAI semantic conventions and tracing agent tool-call workflows. 
  • Data reliability engineering practice: freshness, quality, and lineage SLOs for production data pipelines
  • Telemetry cost engineering or FinOps at scale. 
  • Experience in regulated environments (FedRAMP, GovCloud) and designing telemetry that remains useful under redaction. 
  • Prometheus, Grafana, and Alertmanager. 
  • Snowflake, or experience operating across more than one lakehouse or warehouse platform. 
  • Service catalog or internal developer platform work (Backstage or similar). 

Current US Perks & Benefits:

  • Employer subsidized medical/vision and dental coverage for full-time employees
  • 401k Match to help you save for your future (50% of your contribution up to the first 6% of your eligible pay)
  • Monthly stipend to support your work and productivity
  • Flexible Time Away Program, plus Sick Time Off
  • US employees are automatically covered under Smartsheet-sponsored life insurance, short-term, and long-term disability plans
  • US employees receive 12 paid holidays per year
  • Up to 24 weeks of Parental Leave
  • Personal paid Volunteer Day to support our community
  • Opportunities for professional growth and development including access to Udemy online courses
  • Company Funded Perks, including a counseling membership, local retail discounts, and your own personal Smartsheet account
  • Teleworking options from any registered location in the U.S. (role specific)

Smartsheet provides a competitive base salary range for roles that may be hired in different geographic areas we are licensed to operate our business from. Actual compensation is determined by several factors including, but not limited to, level of professional, educational experience, skills, and specific candidate location. In addition, this role will be eligible for a market competitive incentive opportunity.

US Base Salary Pay Range
$222,500$257,500 USD

 

Get to Know Us:

At Smartsheet, your ideas are heard, your potential is supported, and your contributions have real impact. You’ll have the freedom to explore, push boundaries, and grow beyond your role. We welcome diverse perspectives and nontraditional paths—because we know that impact comes from individuals who care deeply and challenge thoughtfully. When you’re doing work that stretches you, excites you, and connects you to something bigger, that’s magic at work. Let’s build what’s next, together.

Equal Opportunity Employer:

Smartsheet is an Equal Opportunity (EEO) employer committed to fostering an inclusive environment with the best employees. It is our policy to provide equal employment opportunities to all qualified applicants in accordance with applicable laws in the US, UK, Australia, Germany, Costa Rica, Japan, Bulgaria, India, and Singapore. All qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran or disabled status, or genetic information. 

If there are preparations we can make to help ensure you have a comfortable and positive interview experience, please let us know.

 

#LI-Remote

See how well your resume matches this job before you apply

Run a free ATS check