Site Reliability Engineer

Offchain Labs · via Himalayas ·

TypeFull-time job
LocationUnited States
Posted2 hours ago
At Offchain, we aren’t just building products: we’re leading a movement.
As pioneers in blockchain scalability and security, we're at the forefront of transforming how the world interacts with decentralized applications. We're laying the foundation that will define the next generation of digital commerce, governance, and human interaction. This involves tackling real-world challenges that come with scaling blockchain technology, without compromising on its core principles: decentralization, security and transparency.
At the center of this vision is our people. Our team is made up of thinkers and doers that embrace new challenges and seek solutions that push existing boundaries. If you’re energized by solving unprecedented problems, and believe in the role that decentralized systems will play in creating a more equitable digital future, then we want to hear from you.
Why Offchain?
Offchain is setting the pace for the entire Ethereum ecosystem. We built the Arbitrum stack that powers Arbitrum One, the most widely adopted Ethereum scaling solution that exists today.
Arbitrum’s ecosystem is undergoing tremendous growth with hundreds of projects and dApps on Arbitrum One today. Over 100 different teams have used Offchain technology to build their own Arbitrum chains. Major players in the space, Robinhood, BlackRock, Ethena Labs, Securitize, Aave, and Apechain are all using the Arbitrum stack.
Arbitrum’s thriving ecosystem wouldn’t exist without our advanced technology stack. Arbitrum, Prysm, ZeroDev. These aren’t just product names. These are tools that are actively reshaping what's possible on Ethereum and advancing its core infrastructure.
To top it all off? We’re backed by $124 million in funding. We’ve demonstrated consistent execution with billions in secured value, thousands of supported projects, and infrastructure processing millions of transactions seamlessly.
Who You Are

Eager to dive into blockchain technology, even if it’s new territory

Enjoy solving infrastructure problems in unconventional ways and thinking beyond standard patterns

Use tools like k9s or ArgoCD for speed and abstraction, but comfortable dropping into YAML, logs, or low-level debugging when things go sideways

Experienced with GitOps-style systems and treating both infrastructure and application delivery as code

Have scaled deployment automation using patterns like ArgoCD ApplicationSets or similar tooling

Curious about how things work under the hood and not satisfied with surface-level fixes

Comfortable in Linux, fluent in shell scripting, and productive in languages like Python or Go

Comfortable operating within a cloud platform (e.g., AWS, GCP, Azure), with a strong understanding of the underlying components making it easy to adapt to or migrate across providers

Participated in an on-call rotation, responding to incidents, troubleshooting under pressure, and driving postmortems to improve system reliability over time

Design systems with security in mind, applying principles like least privilege and threat modeling

Bring a strong technical foundation, excellent problem-solving skills, and a genuine commitment to high-quality work

Take ownership, collaborate openly, and contribute to a culture of clarity, curiosity, and continuous improvement

What You've Done

Operated production Kubernetes clusters and built scalable, declarative infrastructure using Terraform or similar tools

Deployed and maintained Kubernetes environments, managed system components, and troubleshot applications running on the platform

Designed CI/CD workflows with ArgoCD, GitHub Actions, CodeBuild, or similar tools, covering both infra and app deployments

Designed and operated observability systems using time-series metrics, logs, and dashboards with tools like Prometheus, Loki, Mimir, Grafana, and CloudWatch

Diagnosed tough networking and storage issues across complex, distributed systems

Implemented secure-by-default infrastructure and contributed to architecture reviews and threat models

Automated operational workflows using scripting or programming in Python, Go, or Bash

SREs come from a wide range of backgrounds. If you bring strong problem-solving skills, curiosity, and a drive to build reliable systems, we’d love to hear from you, even if your experience doesn’t perfectly match every bullet point

Perks:

Remote-first global workforce + NY office

Professional reimbursement program (facilitates industry conference attendance, certifications, and more)

Medical, dental & vision coverage (US + some other countries)

401k retirement plan + company match (US only)

Wellness stipend

Home office set up / ergonomic equipment program

Originally posted on Himalayas
site-reliability-engineering engineering devops cloud-engineer blockchain-infrastructure site-reliability-engineer site-reliability-operations-engineer site-reliability-engineering-jobs devops-site-reliability-engineer senior-site-reliability-engineer site-reliability-engineering-(sre) principal-site-reliability-engineer site-reliability-engineering-lead
Apply on Himalayas →

Job sourced from Himalayas. Applications happen directly on the original platform — we never collect your data.