Software Engineer, AI Systems (Canada)

LandingAICanadaRemoteFull-timeStaff, 8–12 yearsListed 1 week ago

Apply now

About this role

What You Will Own:

- Production reasoning systems. Build and operate multi-step LLM pipelines that coordinate model calls, tool calls, graph queries, retrieval, quality gates, and specialist-agent handoffs.

- Agent orchestration. Extend Haven’s coordinated agent team and the orchestration layer that carries an incident from evidence through analysis, review, and enterprise learning.

- Grounding and retrieval. Design the context layer across Neo4j graph traversal, vector search, and hybrid retrieval so every model call receives the right evidence and organizational knowledge.

- Evaluation. Build datasets, scoring, regression suites, model comparisons, human-label loops, and per-stage quality attribution for extraction and reasoning tasks.

- Production AI operations. Implement tracing, tool-call audits, cost and latency monitoring, failure handling, and quality dashboards; catch loops, hallucinations, and silent drift before customers do.

- Technical direction. Select models by task across OpenAI, Anthropic, and Google; partner with product and knowledge engineering; and help shape the AI roadmap.

Requirements For the Role Include:

- Production LLM systems. AI or ML engineering, including shipping LLM systems that real users depend on.

- Agentic workflows. Hands-on experience building and debugging multi-step, tool-calling workflows with LangGraph, LangChain, or an equivalent framework.

- Evaluation discipline. A repeatable approach to LLM evaluation, including representative datasets, regression testing, LLM-as-judge techniques, or human review loops.

- Retrieval judgment. Experience assembling context for LLMs and a clear point of view on what to retrieve, how much, and why.

- Production ownership. A track record of owning systems from deployment through monitoring and incident response, including a strong story about a failure or regression you diagnosed and fixed.

- Model judgment. Comfort working across model providers and explaining tradeoffs in quality, latency, cost, context, and operational risk.

- Graph reasoning. Comfort with Neo4j and Cypher, or a comparable graph store, and the ability to ramp quickly on graph data modeling. Strongly preferred.

- Strong Python. Production habits around FastAPI, asynchronous services, testing, observability, and maintainable interfaces are required.

Nice To Haves Include:

- Deep graph experience. Cypher fluency, schema evolution, MERGE patterns, embeddings, and operating a live knowledge graph.

- Enterprise AI security. Prompt-injection awareness, context-leak prevention, tenant isolation, role-based access, and policy-layer separation.

- Azure and hybrid search. Experience running production AI services in Azure and using Azure AI Search, Pinecone, MongoDB Atlas, pgvector, Elasticsearch, or a similar platform.

- B2B Enterprise SaaS. Prior experience operating in an enterprise-level environment is strongly preferred.

Why Join Haven Safety:

- Meaningful reasoning problems. Work across evidence, causal pathways, controls, organizational history, and corrective actions in a domain where correctness matters.

- An evaluation-first culture. Make quality measurable, observable, and improvable instead of relying on demos or intuition.

- Visible customer impact. Build for safety teams in energy, utilities, infrastructure, construction, and manufacturing, with direct feedback from the people using the output.

- Small team, high ownership. Work closely with the CTO, product, and knowledge engineering, make consequential technical decisions, and see your work reach customers quickly.