AI Engineer — Agentic Systems (TypeScript / Node.js)

Devlight · via Himalayas ·

TypeFull-time job
LocationUkraine
Posted4 hours ago
About us:
Devlight is the AI partner for companies moving from manual work to intelligent operations.
We design and build AI agents, intelligent workflows, and product experiences that reduce manual effort, improve decision-making, and strengthen customer experience.
Built on a decade of delivering digital products used by millions, we bring product discipline to AI, designing systems that teams can operate, trust, and improve over time.
Our clients: Nova Poshta, Sense Bank, Aurora, OKWINE, NOVUS, VARUS, BROCARD, Dnipro-M, and many others.

About the role:
Devlight builds AI systems for enterprise clients, from discovery and architecture through release and production operation.
We are hiring an engineer to build agentic systems. The current engagement is a conversational agent inside a client product with a large user base: the agent interprets the request, decides which tools to call, retrieves data from the client's systems through typed contracts, streams the response into the application, and returns structured actions the application executes itself.
The system is already designed: there is an architecture vision with C1–C3 views, ADRs for every key decision, tool contracts, and specifications for guardrails, grounding, and evaluation methodology. A solution architect owns the architecture and the project tech lead runs day-to-day supervision. Your part is implementation and getting it to release against a written specification.
Technical context: TypeScript / Node.js (NestJS), Anthropic API, streaming to the client, Postgres and Redis, Kubernetes on AWS, end-to-end observability on OpenTelemetry.
This is a project hire of about 3 months, with the scope likely to expand. After successful delivery, the collaboration continues on other projects at the company.

Your future responsibilities:

Build the agent core: the think–act–observe loop, tool selection, session state, layered prompts with caching

Implement the agent's tools and the call runtime: typed contracts, retries, timeouts, error mapping

Guardrails and grounding: input filtering, response validation against retrieved data, safe fallbacks for every failure mode

Streaming to the client: incremental delivery, structured action events for the application, reconnection and backpressure

Evaluation: assemble datasets, build the run harness, wire regression gates before releases and before any model or prompt change

Observability: tracing, cost and latency metrics, dashboards, incident analysis from traces

Extend the stack where the work calls for it: vector and hybrid retrieval, multimodal input, streaming voice

Work from written specifications and keep them current

Your professional qualities:

5+ years of professional experience as a software engineer.

At least one agent you took to production with real users.

Tool design: granularity, argument and response schemas, error and timeout handling, how the agent behaves when a call fails.

Multi-step loops and context: layered prompt structure with caching, session and conversation state, parallel tool calls, recovery after failure, token and latency budgets.

LLM APIs at the level of detail: tool use, structured outputs, streaming, prompt caching, limits and retries.

Grounding and guardrails: validating responses against the source, input filtering, off-topic and jailbreak handling, discipline around personal data.

Choosing the grounding approach for the data at hand: direct tool calls or RAG, with the decision argued.

Evaluation as a practice: test case sets, automated grading with LLM-as-judge, regression gates before a release and before any prompt or model change.

Tracing and metrics: every model call and tool call as a span, cost and latency per conversation, diagnosing an incident from the trace.

Production backend work in TypeScript / Node.js: NestJS or a comparable framework, REST, streaming over SSE or WebSocket, Postgres and Redis.

Claude Code (or a comparable agentic coding tool) in daily use: subagents, skills, hooks, MCP, defining repository conventions.

Spec-driven development: specifications and contracts come before code, decisions are captured in ADRs, documentation is updated alongside changes.

Nice to have:

Full-stack: internal interfaces in React / TypeScript (metrics dashboards, conversation browsing, prompt version management) — with this, one person covers the whole role.

Deeper RAG mechanics: chunking and metadata, embedding model selection, hybrid retrieval, reranking, retrieval metrics (recall@k, MRR / nDCG).

Search in production: Elasticsearch or OpenSearch at the relevance-tuning level, pgvector.

Multimodal models and streaming speech-to-text: CLIP / SigLIP class embeddings, vision models, latency budgeting for a voice pipeline.

Evaluation and observability tooling: Promptfoo, LangSmith, Langfuse, Braintrust, OpenTelemetry, Prometheus, Grafana, Loki.

Infrastructure and data: Kubernetes, Terraform, queues, BigQuery and analytical pipelines.

Products with a large user base, where latency and cost per interaction matter.

What we offer for your success:

Fully remote or hybrid work format.

Paid Time Off, sick days, regular reward evaluations, and accounting support.

Reimbursement for training courses and compensation for the use of personal equipment.

IT Club Loyalty Card.

Work with an open-minded team that welcomes your new ideas, alongside the best specialists who love sharing their experience.

Get the chance to connect with top companies and contribute to the growth of the Ukrainian IT community together.

Our recruitment process:
Recruiter Interview ✅ Tech Interview ✅ Reference Check ✅ Offer ✅
Ready to become a part of Devlight? Go ahead and send us your CV. We’ll be thrilled to welcome you to the team!Originally posted on Himalayas
ai-engineering backend-engineering software-engineer agentic-systems full-stack-engineering senior-agentic-ai-engineer senior-agentic-systems-developer ai-engineer
Apply on Himalayas →

Job sourced from Himalayas. Applications happen directly on the original platform — we never collect your data.