OA
Site Reliability Engineer
Accepting applicationsOtto AI · Bengaluru, Karnataka, India
Contract Entry AISOC
Estimated market salary
₹6-11 LPA
This is a SiliconBoard market estimate, not an employer-posted salary.
Posted
1d ago
Category
Manufacturing
Experience
Entry
Country
India
AWS · TypeScript · Reliability · Security · Scale
Reliability, scale, security, and cost are all yours.
Otto builds an AI computer.
Our hardware sits on a customer's desk running embedded Linux, maintains a live connection to our cloud, and receives signed over-the-air updates from us.
Every part of that path runs on AWS. Our infrastructure is infrastructure-as-code, and our stack is heavily TypeScript.
You will design and operate the platform behind Otto: cloud infrastructure, CI/CD, observability, fleet reliability, security, capacity planning, and cost.
You will also write code, review PRs, deploy to production, respond to incidents, and participate in on-call.
This is a small team. There is no infrastructure layer between you and the system.
Infrastructure as Code
AWS CDK v2
TypeScript
VPC
IAM
Compute
AWS Lambda
ECS Fargate
API Gateway
CloudFront
Data
RDS PostgreSQL
DynamoDB
ElastiCache / Redis
S3
Delivery
GitHub Actions
Docker
ECR
Canary, rolling, and blue/green deployments
Automated rollback
Observability
CloudWatch
OpenTelemetry
Grafana
SLIs / SLOs
Alerting and incident response
Nearby Stack
Node.js
Next.js
Vercel
Some of our web surfaces run on Vercel. Experience with it is helpful, but not required.
What You'll OwnProduction Infrastructure on AWS
Design, build, and operate our production infrastructure across Lambda, Fargate, RDS, ElastiCache, DynamoDB, S3, CloudFront, API Gateway, VPC, and IAM.
You should be comfortable owning the system end to end rather than relying on a separate platform team.
Infrastructure as Code
Every resource lives in our CDK tree.
Development and production environments should be reproducible from the same code, with infrastructure changes reviewed through the same PR process as application code.
CI/CD
Build and maintain GitHub Actions pipelines for:
Application deployments
Container deployments
Infrastructure deployments
Database migrations
Canary releases
Rolling deployments
Blue/green deployments
Automated rollback
You should know when each deployment strategy is appropriate and why.
Observability
Own metrics, logs, traces, dashboards, and alerts across the platform.
The goal is simple: page a human only when a human is actually needed.
SLIs, SLOs, and Error Budgets
Define reliability targets, track them, and help the product and engineering teams make informed decisions when reliability budgets are being consumed.
Incident Response
Lead incidents when they happen.
Run blameless postmortems and turn findings into real reliability work rather than documentation that gets forgotten.
OTA Updates to the Otto Fleet
Our releases eventually reach physical devices sitting in customers' homes and offices.
You will help own:
Canary rollouts
Rollback detection
Fleet health monitoring
Release safety
Failure recovery
A bad release can reach hardware we cannot physically access, so deployment discipline matters.
AWS Cost
Own infrastructure efficiency and visibility.
This includes:
Rightsizing
Spot and reserved capacity
Tagging discipline
Budget alerts
Cost attribution
Identifying which services and workloads are actually driving spend
Security
Help establish and maintain a strong production security posture, including:
IAM least privilege
Network segmentation
Secrets management
Patching
Access controls
Auditability
Foundations for SOC 2
Data Infrastructure
Operate PostgreSQL, Redis, and DynamoDB in production.
That includes:
Backups
Tested restores
Capacity planning
Upgrades
Performance
Availability
Failure recovery
Killing Toil
If something repetitive can be automated, automate it.
Build internal tooling in TypeScript or Bash that allows engineers to self-serve instead of relying on manual infrastructure work.
Mentoring
Help raise the operational and reliability bar across the entire engineering team.
Must Have
These are not "familiarity with" requirements.
We are looking for someone who has owned these systems in production and understands where they fail.
AWS CDK
Deep, current experience with AWS CDK v2 in TypeScript, including:
Constructs
Stacks
Cross-stack references
Cross-region architecture
Custom resources
Deployment troubleshooting
Managing infrastructure changes safely in production
You should know what to do when the synth is clean but the deployment still isn't.
AWS in Production
Strong hands-on experience with:
VPC architecture and networking
IAM policy design
CloudFront
API Gateway
S3
Production security and access controls
We are looking for direct ownership, not experience where another platform team handled the difficult parts.
AWS Lambda
You should understand:
Cold starts
Concurrency limits
VPC-attached functions
Bundle size
Scaling behavior
Timeouts
Connection management
The problems Lambda creates when talking to PostgreSQL
RDS PostgreSQL
Experience actually operating PostgreSQL in production, including:
Query plans
Indexing
Vacuum behavior
Connection limits
RDS Proxy
Major-version upgrades
Point-in-time recovery
Backup and restore procedures you have actually tested
DynamoDB
Strong understanding of DynamoDB data modeling, including:
Designing around access patterns
Partition keys
GSIs
Hot partitions
Conditional writes
On-demand vs. provisioned capacity
TypeScript and Node.js
You should write real production TypeScript, not just infrastructure glue.
Our infrastructure, backend, internal tooling, and device runtime are all heavily Node.js and TypeScript.
You will regularly read and review code outside of a traditional infrastructure role.
Linux and Networking
Strong understanding of:
TCP/IP
DNS
TLS
Load balancing
WebSockets
Linux systems
You should be comfortable debugging problems like a persistent socket that only drops for customers on one ISP.
CI/CD and Incident Response
You should have experience:
Building deployment pipelines
Choosing deployment strategies
Running production incidents
Writing postmortems
Establishing SLOs
Turning incidents into reliability improvements
Bonus
These are not required, but they count for a lot.
Next.js, including App Router, Server Components, and build pipelines
Vercel, including projects, preview environments, edge configuration, and hosted infrastructure
Edge, IoT, or embedded Linux fleets with OTA updates
Redis pub/sub
Large-scale WebSocket services
AWS Organizations and multi-account architectures
OIDC deployment roles
Ephemeral per-developer environments
LLM gateways, proxies, or AI workloads on AWS
SOC 2 or ISO 27001 experience
FinOps experience
Self-hosted CI runners
Container build caching
Kubernetes experience — useful context, although we do not currently run Kubernetes
AWS certifications such as Solutions Architect, DevOps Engineer, or SysOps Administrator
What This Is Really Like
Otto is a small team with a short path from decision to production.
You will work in the same repositories as the rest of the engineering organization and regularly read code outside your lane.
Our systems span cloud infrastructure, APIs, data infrastructure, web applications, and physical devices running in the field.
We value clear technical writing.
Design documents, postmortems, architecture decisions, and the sentence in a PR explaining why something is changing all matter.
If a deployment changes something persisted on an Otto device that has been running in someone's home for six months, we want that risk understood and documented before the release goes out.
On-call is shared and real.
In return, fixing the system that woke you up is treated as the work — not as an interruption from the work.
Show more Show less
Reliability, scale, security, and cost are all yours.
Otto builds an AI computer.
Our hardware sits on a customer's desk running embedded Linux, maintains a live connection to our cloud, and receives signed over-the-air updates from us.
Every part of that path runs on AWS. Our infrastructure is infrastructure-as-code, and our stack is heavily TypeScript.
You will design and operate the platform behind Otto: cloud infrastructure, CI/CD, observability, fleet reliability, security, capacity planning, and cost.
You will also write code, review PRs, deploy to production, respond to incidents, and participate in on-call.
This is a small team. There is no infrastructure layer between you and the system.
Infrastructure as Code
AWS CDK v2
TypeScript
VPC
IAM
Compute
AWS Lambda
ECS Fargate
API Gateway
CloudFront
Data
RDS PostgreSQL
DynamoDB
ElastiCache / Redis
S3
Delivery
GitHub Actions
Docker
ECR
Canary, rolling, and blue/green deployments
Automated rollback
Observability
CloudWatch
OpenTelemetry
Grafana
SLIs / SLOs
Alerting and incident response
Nearby Stack
Node.js
Next.js
Vercel
Some of our web surfaces run on Vercel. Experience with it is helpful, but not required.
What You'll OwnProduction Infrastructure on AWS
Design, build, and operate our production infrastructure across Lambda, Fargate, RDS, ElastiCache, DynamoDB, S3, CloudFront, API Gateway, VPC, and IAM.
You should be comfortable owning the system end to end rather than relying on a separate platform team.
Infrastructure as Code
Every resource lives in our CDK tree.
Development and production environments should be reproducible from the same code, with infrastructure changes reviewed through the same PR process as application code.
CI/CD
Build and maintain GitHub Actions pipelines for:
Application deployments
Container deployments
Infrastructure deployments
Database migrations
Canary releases
Rolling deployments
Blue/green deployments
Automated rollback
You should know when each deployment strategy is appropriate and why.
Observability
Own metrics, logs, traces, dashboards, and alerts across the platform.
The goal is simple: page a human only when a human is actually needed.
SLIs, SLOs, and Error Budgets
Define reliability targets, track them, and help the product and engineering teams make informed decisions when reliability budgets are being consumed.
Incident Response
Lead incidents when they happen.
Run blameless postmortems and turn findings into real reliability work rather than documentation that gets forgotten.
OTA Updates to the Otto Fleet
Our releases eventually reach physical devices sitting in customers' homes and offices.
You will help own:
Canary rollouts
Rollback detection
Fleet health monitoring
Release safety
Failure recovery
A bad release can reach hardware we cannot physically access, so deployment discipline matters.
AWS Cost
Own infrastructure efficiency and visibility.
This includes:
Rightsizing
Spot and reserved capacity
Tagging discipline
Budget alerts
Cost attribution
Identifying which services and workloads are actually driving spend
Security
Help establish and maintain a strong production security posture, including:
IAM least privilege
Network segmentation
Secrets management
Patching
Access controls
Auditability
Foundations for SOC 2
Data Infrastructure
Operate PostgreSQL, Redis, and DynamoDB in production.
That includes:
Backups
Tested restores
Capacity planning
Upgrades
Performance
Availability
Failure recovery
Killing Toil
If something repetitive can be automated, automate it.
Build internal tooling in TypeScript or Bash that allows engineers to self-serve instead of relying on manual infrastructure work.
Mentoring
Help raise the operational and reliability bar across the entire engineering team.
Must Have
These are not "familiarity with" requirements.
We are looking for someone who has owned these systems in production and understands where they fail.
AWS CDK
Deep, current experience with AWS CDK v2 in TypeScript, including:
Constructs
Stacks
Cross-stack references
Cross-region architecture
Custom resources
Deployment troubleshooting
Managing infrastructure changes safely in production
You should know what to do when the synth is clean but the deployment still isn't.
AWS in Production
Strong hands-on experience with:
VPC architecture and networking
IAM policy design
CloudFront
API Gateway
S3
Production security and access controls
We are looking for direct ownership, not experience where another platform team handled the difficult parts.
AWS Lambda
You should understand:
Cold starts
Concurrency limits
VPC-attached functions
Bundle size
Scaling behavior
Timeouts
Connection management
The problems Lambda creates when talking to PostgreSQL
RDS PostgreSQL
Experience actually operating PostgreSQL in production, including:
Query plans
Indexing
Vacuum behavior
Connection limits
RDS Proxy
Major-version upgrades
Point-in-time recovery
Backup and restore procedures you have actually tested
DynamoDB
Strong understanding of DynamoDB data modeling, including:
Designing around access patterns
Partition keys
GSIs
Hot partitions
Conditional writes
On-demand vs. provisioned capacity
TypeScript and Node.js
You should write real production TypeScript, not just infrastructure glue.
Our infrastructure, backend, internal tooling, and device runtime are all heavily Node.js and TypeScript.
You will regularly read and review code outside of a traditional infrastructure role.
Linux and Networking
Strong understanding of:
TCP/IP
DNS
TLS
Load balancing
WebSockets
Linux systems
You should be comfortable debugging problems like a persistent socket that only drops for customers on one ISP.
CI/CD and Incident Response
You should have experience:
Building deployment pipelines
Choosing deployment strategies
Running production incidents
Writing postmortems
Establishing SLOs
Turning incidents into reliability improvements
Bonus
These are not required, but they count for a lot.
Next.js, including App Router, Server Components, and build pipelines
Vercel, including projects, preview environments, edge configuration, and hosted infrastructure
Edge, IoT, or embedded Linux fleets with OTA updates
Redis pub/sub
Large-scale WebSocket services
AWS Organizations and multi-account architectures
OIDC deployment roles
Ephemeral per-developer environments
LLM gateways, proxies, or AI workloads on AWS
SOC 2 or ISO 27001 experience
FinOps experience
Self-hosted CI runners
Container build caching
Kubernetes experience — useful context, although we do not currently run Kubernetes
AWS certifications such as Solutions Architect, DevOps Engineer, or SysOps Administrator
What This Is Really Like
Otto is a small team with a short path from decision to production.
You will work in the same repositories as the rest of the engineering organization and regularly read code outside your lane.
Our systems span cloud infrastructure, APIs, data infrastructure, web applications, and physical devices running in the field.
We value clear technical writing.
Design documents, postmortems, architecture decisions, and the sentence in a PR explaining why something is changing all matter.
If a deployment changes something persisted on an Otto device that has been running in someone's home for six months, we want that risk understood and documented before the release goes out.
On-call is shared and real.
In return, fixing the system that woke you up is treated as the work — not as an interruption from the work.
Show more Show less
Similar Jobs
AT
Photonic Design & Fab Engineer
Acceler8 Talent · Boulder, CO
KT
Hardware, Optical and Electrical Engineer
Keysight Technologies · Santa Rosa, CA
S
Controls Engineer, Supercomputer Infrastructure - Memphis
SpaceXAI · Southaven, MS
TT
Embedded Firmware Engineer - The Toro Company
The Toro Company · Riverside, CA