About this role
Responsibilities
- Own critical subsystems and large features end to end - from problem discovery and technical design through development, rollout, monitoring, and measuring customer and business impact.
- Lead major technical initiatives within the team and drive architecture, design, implementation, rollout strategy, and post-release improvements.
- Proactively identify system problems, including scalability bottlenecks, reliability risks, performance issues, technical debt, usability gaps, and operational inefficiencies, and drive solutions without waiting for them to be assigned.
- Design and operate reliable distributed systems using message queues, background workers, schedulers, retries, idempotency, caching, and high-volume third-party platform API integrations.
- Anticipate scale, edge cases, failure modes, rate limits, API changes, and downstream system implications before they become production problems.
- Evaluate architectural and technology trade-offs across scalability, reliability, maintainability, cost, development velocity, and customer impact.
- Build data pipelines that ingest and enrich large volumes of social data and metrics from multiple sources and provide infrastructure for fast analytics and insights.
- Process media at scale, including image and video upload, transformation, optimization, generation, storage, and delivery in the formats required by different social platforms.
- Ship both backend APIs and frontend experiences, with strong attention to usability, performance, accessibility, and customer experience.
- Build production-grade AI capabilities using LLMs, image/video generation, embeddings, RAG, tool calling, and agentic workflows to improve customer outcomes.
- Identify opportunities where AI and intelligent technologies can improve product capabilities, engineering productivity, system reliability, or operational efficiency, and validate them through prototypes and measurable outcomes.
- Follow and promote Test-Driven Development and strong automated testing practices to reduce defect leakage and improve release confidence.
- Own production quality through observability, monitoring, incident investigation, root-cause analysis, and permanent corrective actions.
- Raise the engineering quality bar through code reviews, design reviews, engineering standards, documentation, testing practices, and maintainable architecture.
- Mentor engineers and help the team make better technical decisions through design discussions, reviews, prototypes, and knowledge sharing.
Requirements
- 4+ years of hands-on software engineering experience building and operating production systems, preferably SaaS products at meaningful scale.
- Strong experience independently owning complex or business-critical features or services from design through production.
- Hands-on experience building distributed systems using message queues, background workers, schedulers, retries, idempotency, caching, and high-volume third-party API integrations.
- Extensive experience integrating with third-party APIs and handling real-world challenges such as rate limits, authentication and token management, webhooks, retries, API version changes, partial failures, and inconsistent external dependencies.
- Strong ability to proactively identify system problems and drive improvements across scalability, reliability, performance, maintainability, and technical debt.
- Strong understanding of system design and the ability to evaluate architectural trade-offs, edge cases, downstream dependencies, and long-term implications.
- Strong engineering fundamentals with experience writing clean, maintainable, well-tested code and applying Test-Driven Development or similar disciplined testing practices.
- Strong TypeScript and Node.js experience; hands-on experience with NestJS is preferred.
- Experience building customer-facing web applications using Vue.js, React, Angular, or similar modern frontend frameworks, with a good understanding of UI/UX principles and the ability to identify usability gaps.
- Experience with MongoDB, Redis, and data-intensive applications.
- Experience with large-scale data ingestion, ETL/ELT, analytics, or data warehousing using technologies such as ClickHouse, BigQuery, Snowflake, Elasticsearch, or similar systems.
- Hands-on experience building AI-powered software using LLM APIs, embeddings, RAG, image/video generation, tool calling, or agentic workflows.
- Ability to move AI features beyond prototypes into reliable production systems with appropriate evaluation, observability, guardrails, cost considerations, and failure handling.
- Hands-on experience with AI-assisted software development and the ability to use AI coding tools thoughtfully while maintaining ownership of architecture, quality, security, and correctness.
- Customer-first and product-oriented mindset: start with the customer or business problem and select technology based on the desired outcome.
- Strong written and verbal communication skills, with the ability to lead technical discussions, communicate trade-offs clearly, and collaborate effectively in a remote environment.
