About this role
company
About Luma Financial Technologies
Founded in 2018, Luma Financial Technologies (“Luma”) has pioneered a cutting-edge fintech software platform that has been adopted by broker/dealer firms, RIA offices, and private banks around the world. By using Luma, institutional and retail investors have a fully customizable, independent, buy-side technology platform that helps financial teams more efficiently learn about, research, purchase, and manage alternative investments as well as annuities. Luma gives these users the ability to oversee the full, end-to-end process lifecycle by offering a suite of solutions. These include education resources and training materials; creation and pricing of custom structured products; electronic order entry; and post-trade management. By prioritizing transparency and ease of use, Luma is a multi-issuer, multi-wholesaler, and multi-product option that advisors can utilize to best meet their clients’ specific portfolio needs. Headquartered in Cincinnati, OH, Luma also has offices in New York, NY, Miami, FL, Zurich, Switzerland and Lisbon, Portugal. For more information, please visit Luma’s website .
role
About the role
We are seeking a proactive and skilled Production Support Engineer to join our IT operations team. In this critical role, you will be responsible for overseeing and managing incidents related to our engineering and production environments. The successful candidate will play a key role in ensuring the stability and reliability of our production systems by coordinating and resolving incidents in a timely and efficient manner, working closely with Software Engineers, Infrastructure, and Support teams to triage, diagnose, and resolve technical issues in real time
What you'll do
Application Monitoring & Incident Management
- L2/L3 Incident Response: Triage, investigate, and resolve production bugs, system outages, and customer-impacting issues within SLA guidelines.
- Proactive Monitoring: Monitor engineering and production system performance, application metrics, and server logs using tools like Datadog.
- Root Cause Analysis (RCA): Lead post-incident write-ups and collaborate with developers to design long-term fixes to prevent recurring incidents.
Technical Operations & Maintenance
- Deployment & Releases: Assist with blue-green deployments, hotfixes, and scheduled production maintenance windows.
- Environment Management: Maintain high availability across production, staging, and disaster-recovery (DR) environments.
- Database & Data Fixes: Execute safely managed SQL queries, data migrations, or manual record reconciliations when required.
Automation & Process Improvement
- Automation: Write scripts (Bash, Python, PowerShell) to automate repetitive operational tasks, log parsing, and monitoring alerts.
- Runbook Documentation: Create, update, and standardize operational runbooks, escalation matrixes, and troubleshooting guides.
Qualifications
Required
- Education : Bachelor’s degree in Computer Science, Information Technology, or a related discipline (or equivalent practical experience).
- Experience: 3+ years in a Production Support, Site Reliability Engineering (SRE), or Application Support role.
- Operating Systems: Solid command-line skills in Linux/Unix environments.
- Database / SQL: Strong ability to write complex SQL queries , analyze slow queries, and understand database schemas (MongoDB, MySQL, or MS SQL).
- Scripting: Proficiency in at least one scripting language ( Python, Bash, Shell, or PowerShell ) for automation.
- Monitoring Tools: Hands-on experience with log aggregation and monitoring platforms (e.g., Datadog).
- Protocols & Architecture: Understanding of REST APIs, microservices architecture, network fundamentals (TCP/IP, DNS, HTTP statuses), and web servers (Nginx, Apache).
- Soft Skills : Excellent analytical skills to troubleshoot complex system failures under pressure, alongside the communication skills required to translate technical ideas to non-technical stakeholders.
- On-Call Availability: Willingness to participate in a rotational on-call schedule (24/7 support coverage).
Preferred
- Platform experience : Experience with cloud platforms ( AWS or Azure ). Familiarity with containerization and orchestration ( Kubernetes ). Experience with ticketing systems like Jira. Exposure to CI/CD tools (GitHub Actions, GitLab CI).
- Industry Exposure : Prior Exposure to fintech, insurance or broader finance services organizations.