About this role
About Skan AI Be at the Forefront of the Agentic AI Revolution At Skan AI, we are pioneering the context engine for human and agentic execution, bringing context from enterprise operators, systems, and processes to power how the world's largest organizations execute their most complex, mission-critical work.
Why Skan AI We're in hyper-growth mode at exactly the right moment in history. As enterprises race to adopt agentic AI, we're uniquely positioned to deliver the clear signal they desperately need: a platform that trains and grounds AI Agents in trillions of real execution signals, enabling reliable, compliant automation of their most complex processes.
Backed by Dell Technologies Capital and other leading investors, we're the only company that can bridge the gap between AI's promise and enterprise reality, making us perfectly positioned to define the agentic era for modern enterprises.
Our diverse, collaborative team of 250+ innovators is solving category-defining challenges at the intersection of AI, process intelligence, and enterprise work. Diverse perspectives fuel breakthrough thinking, cross-functional collaboration is the norm, and our work directly transforms how Fortune 500 companies operate. We are shaping the future of work itself.
# About the Role
Skan is looking for a Senior Architect Leader to build and lead its L2 function: the team that owns the deepest, most complex technical issues our customers face on Skan platforms. These issues span data quality and integrity, network connectivity, virtual assistants, API integrations, data correlation, application performance, data management, and more. This is a hands-on leadership role. You will lead a team of senior engineers, personally dive into the hardest problems, and act as the technical bridge between customers and Product Engineering. You will be the person who connects the dots across seemingly unrelated events, finds the true root cause, and makes sure the same problem does not happen twice. We are looking for someone who is highly responsive, collaborative, and energetic, and who brings deep experience across application architecture, databases, data models, LLMs, API integration, and network connectivity.
Key Responsibilities
## Leadership of the L2 Function
Build, lead, mentor, and grow a high-performing team of L2 engineers, setting clear expectations around ownership, responsiveness, and technical rigor. Define and continuously improve the L2 operating model, including escalation paths from L1, handoffs to Product Engineering, triage priorities, SLAs, and on-call coverage for critical issues. Track and report on key metrics such as time to resolution, escalation rates, recurrence rates, customer satisfaction, and backlog health, and use them to drive improvement.
## Deep Technical Troubleshooting
Act as the senior-most technical escalation point for complex customer issues across the Skan platform, including:
- Data quality and integrity: missing, duplicated, inconsistent, or corrupted data; data pipeline and ingestion failures.
- Networking and connectivity: proxy, firewall, DNS, TLS/SSL, VPN, and endpoint-to-cloud connectivity problems in customer environments.
- Virtual assistants and LLM-powered features: unexpected responses, accuracy or grounding issues, latency, prompt and context behavior, and model integration failures.
- API integrations: authentication and authorization failures, rate limiting, payload and schema mismatches, and third-party system integration issues.
- Data correlation: issues in how events, sessions, users, and activities are stitched together and interpreted across sources.
- Application performance: slowness, resource contention, scaling bottlenecks, memory and CPU issues, and query performance.
- Data management: storage, retention, migration, backup and restore, and configuration issues.
Analyze logs, telemetry, traces, metrics, database records, and configuration data to isolate problems quickly and accurately. Correlate multiple events across components, time windows, and customer environments to form and validate hypotheses and arrive at sound conclusions. Reproduce issues in controlled environments when needed, and provide clear workarounds to customers while permanent fixes are developed.
## Collaboration with Product Engineering
Partner with Product Engineering to resolve critical bugs and production issues, providing well-documented, reproducible, and prioritized escalations. Participate in and, where appropriate, lead incident war rooms for high-severity issues, coordinating across engineering, operations, and customer-facing teams. Feed field insights back into the product roadmap, highlighting supportability gaps, recurring defects, and opportunities to improve observability, diagnostics, and resilience.
## Root Cause Analysis and Prevention
Lead detailed Root Cause Analyses (RCAs) for significant and recurring issues, clearly documenting what happened, why it happened, its impact, and the corrective and preventive actions required. Drive proactive measures to prevent recurrence, such as monitoring and alerting improvements, configuration hardening, diagnostic tooling, automated health checks, and product fixes. Identify trends and patterns across tickets to get ahead of issues before customers report them.
## Technical Decision-Making and Consensus Building
Drive technical decisions on complex and ambiguous issues by building consensus across engineering, product, support, and customer stakeholders. Present findings and recommendations clearly, backed by data, and navigate disagreements constructively toward the best outcome for the customer and the platform.
## Knowledge Management
Lead the documentation of recurring issues and their fixes as high-quality Knowledge Base (KB) articles, runbooks, and troubleshooting guides. Establish standards and review processes for KB content to keep it accurate, searchable, and current. Use the knowledge base to enable L1 to resolve more issues independently and reduce escalations to L2.
## Customer Engagement
Communicate directly with customer technical teams (IT, security, infrastructure, and data teams) during critical escalations, translating complex technical findings into clear explanations and action plans. Build trust with key customers through responsiveness, transparency, and follow-through.
Required Qualifications
- 12+ years of experience in software engineering, solutions architecture, technical support engineering, or site reliability engineering, including at least 4 years leading technical teams.
- Deep experience in application architecture, including distributed systems, microservices, cloud-native platforms, and enterprise SaaS.
- Strong expertise in databases and data models, including SQL and NoSQL systems, query optimization, schema design, and data pipeline troubleshooting.
- Hands-on experience with LLMs and AI-powered applications, including how virtual assistants and generative AI features are built, integrated, and debugged.
- Solid understanding of API design and integration, including REST, authentication standards (OAuth, SAML, API keys, tokens), webhooks, and enterprise system integrations.
- Strong knowledge of network connectivity and security, including TCP/IP, DNS, HTTP/S, TLS, proxies, firewalls, load balancers, and enterprise network constraints.
- Proven ability to troubleshoot complex, multi-layer issues and to connect the dots across multiple events, systems, and data sources to reach well-supported conclusions.
- Proficiency with log analysis, observability, and monitoring tools (for example Splunk, Datadog, ELK, Grafana, New Relic, or similar).
- Demonstrated track record of conducting rigorous RCAs and driving preventive actions to completion.
- Experience writing and maintaining technical documentation, KB articles, and runbooks.
- Excellent communication skills, with the ability to engage confidently with both engineers and customer executives.
Preferred Qualifications
- Experience with process intelligence, process mining, analytics, or enterprise monitoring platforms.
- Experience with endpoint or desktop agents deployed in enterprise environments.
- Familiarity with major cloud platforms (AWS, Azure, or GCP), containers, and Kubernetes.
- Scripting skills (Python, PowerShell, Bash, or similar) for diagnostics and automation.
- Experience with ITSM and support tooling (for example Jira, ServiceNow, Zendesk, or Freshdesk).
- Exposure to ITIL practices, incident management, and problem management frameworks.
- Experience building a support or escalation function from the ground up in a high-growth company.
Personal Attributes
- Highly responsive: acts with urgency when customers are impacted and keeps stakeholders informed.
- Collaborative: works across teams with humility and builds strong partnerships with engineering and customer-facing groups.
- Energetic and driven: brings positive energy, ownership, and persistence to difficult problems.
- Analytical and curious: never settles for the symptom and keeps digging until the true root cause is found.
- Consensus builder: influences without authority and aligns diverse stakeholders around sound technical decisions.
- Calm under pressure: stays composed and structured during high-severity incidents.
- Teacher and mentor: raises the technical bar of the team and shares knowledge generously.
Skan AI is an equal opportunity employer committed to building a diverse, inclusive, and respectful workplace around the world. We do not discriminate based on race, color, religion or belief, sex (including pregnancy, sexual orientation, gender identity, or gender expression), national origin, ancestry, age, disability, medical condition, genetic information, marital or family status, military or veteran status, or any other characteristic protected by applicable laws in the locations where we operate. We welcome people from all backgrounds and provide reasonable accommodations throughout the hiring process.