About this role
company
Aerospike is the real-time database for mission-critical use cases and workloads, including machine learning, generative, and agentic AI. Aerospike powers millions of transactions per second with millisecond latency, at a fraction of the total cost of ownership compared to other databases.
Global leaders, including Adobe, Airtel, Barclays, Criteo, DBS Bank, Experian, Grab, HDFC Bank, PayPal, Sony Interactive Entertainment, The Trade Desk, and Wayfair, rely on Aerospike for customer 360, fraud detection, real-time bidding, profile stores, recommendation engines, and other use cases.
At Aerospike, we dream big and deliver even bigger. Our mission is to unleash the power of the world’s real-time data with a database built for infinite scale, speed, and sustainability .
If you're ready to shape the future of data, join us.
role
About the role
We are hiring a Principal Software Quality Engineer to own the quality bar for a distributed database product (NoSQL preferred). This is a senior individual-contributor role for someone who lives inside the system under test — reads the server code, reproduces bugs down to the offending function, designs fault-injection to exercise the paths that positive tests can’t reach, and drives fixes across QE, client, and server teams.
You will also lead how AI agents show up inside our testing frameworks: agents that generate and maintain tests, triage failures against logs, traces, metrics, and database state, and are themselves evaluated for reliability.
What you will do
- Own the quality strategy for a distributed database platform: test architecture, coverage model, release gates, and risk-based prioritization across server, client SDKs, operator/backup tooling, and management surfaces.
- Design and drive deep protocol- and storage-level testing : replication and rebalancing, partitioning/sharding, consistency (strong, eventual, tunable), transactions and isolation, secondary indexes, query planning, compaction, backup and restore, snapshots, warm/cold restart, and cross-datacenter replication.
- Prove correctness of the cluster under stress: node failures, network partitions, clock skew, disk-full and slow-disk conditions, rolling and mixed-version upgrades, and long-running soak.
- Use fault injection — process crashes, gdb-driven state manipulation, packet drops, latency injection, corrupted wire data, oversized/boundary inputs, and forced version skew between master and replicas — to reach code paths positive tests can’t.
- Perform root-cause investigations that reach into the source : identify off-by-one errors, out-of-bounds reads, encoder/decoder mismatches, compression/CRC issues, and consistency violations; hand dev a minimal repro plus the suspect function and file crisp tickets that drive quick fixes.
- Build and evolve high-leverage test frameworks (Python-based pytest / harness code, plus API and end-to-end suites) that run reliably on a shared test farm and in CI (GitHub Actions,Jenkins, or equivalent).
- Lead AI-agent testing : design agent loops that plan tests, call tools, inspect logs/traces/metrics/DB state, and produce actionable defects; build evaluation harnesses (golden sets, graders, regression suites) so agent output is trusted, not guessed. Integrate AI agents into existing frameworks (pytest, custom harnesses, MCP-style tool servers) so they can generate negative and corner-case scenarios, maintain flaky tests, and triage failures with a human in the loop where it matters.
- Represent quality in feature reviews (PRDs, design docs, client specs): author test plans, drive dev sessions on new features, and land testability requirements before code is written. Turn production and pre-release incidents into durable tests, regressions, and monitors; keep flake and noise on the test farm low.
- Mentor senior QEs and SDETs on debugging, test design, and framework code; raise the org-wide quality bar without needing a management chain.
What you bring
Required
- 10 - 18 years in software quality / SDET / systems engineering with a quality-systems focus; demonstrated Principal or Staff-level impact.
- Deep, hands-on testing of a distributed database or storage system — NoSQL preferred (e.g. Aerospike, Cassandra, ScyllaDB, MongoDB, Couchbase, Redis, DynamoDB, HBase) — or a relational distributed store (CockroachDB, TiDB, YugabyteDB, Spanner-style). You have tested replication, partitioning/sharding, consistency, transactions, secondary indexes, backup/restore, and cross-datacenter replication — not just applications on top of a database.
- Solid grounding in distributed-systems fundamentals: replication protocols, consensus, quorum reads/writes, eventual vs strong consistency, split-brain, clock skew, and failure/recovery semantics.
- Ability to read server code (C, C++, Go, Rust, or Java) well enough to localize a bug to a function and explain root cause to dev, including protocol/serialization edge cases (wire formats, msgpack/protobuf/Thrift, compression codecs).
- Strong software engineering skills in Python (primary) and comfort in at least one systems language; you treat test infrastructure as a product, not a script pile.
- Practical fault-injection experience : gdb, chaos-style tooling, kill/quiesce, network partitions, latency/packet-loss injection, disk faults, corrupted inputs, forced version skew, and boundary/oversized data.
- Multi-layer testing fluency: unit, integration, API, contract, end-to-end, performance/scale, upgrade/downgrade, and long-running soak/regression.
- Proven experience running debug threads across QE, client, and server teams: you write the analysis, propose the fix location, and follow up until master/dev/stage branches carry the patch.
- Solid CI/CD background (GitHub Actions, Jenkins, or equivalent), and comfort operating a shared test farm: registering suites, triaging noisy failures, and driving flake down. Practical AI-agent-for-testing experience : agent loops, tool use, prompt/eval design, and using agents against real test frameworks. You can tell an agent doing useful work from one producing plausible-looking noise.
- Excellent written communication: your Slack threads, tickets, and RFCs are the kind other engineers save and quote.
Strongly preferred
- Experience testing client SDKs in multiple languages (Python, Java, Go, C, Node.js, .NET) against the same server, including compatibility matrices, upgrade/downgrade, and behavioral parity.
- Backup/restore, checkpointing, snapshotting, or WAL/journal correctness testing, including performance at TB-scale.
- Query-engine or secondary-index testing: query planning, index selection, pagination, boundary values, and encoding/collation correctness.
- Kubernetes / operator testing (KUTTL, Chainsaw, envtest) for database operators. Cloud infrastructure fluency for large-scale test environments (AWS/GCP/Azure, TLS/mTLS, IAM, secrets, air-gapped installs).
- Security-adjacent testing: fuzzing, CVE reproduction, memory-safety issues (OOB reads, use-after-free) in native code.
How we will measure success
- Critical database paths — replication, rebalancing, consistency, transactions, secondary indexes, backup/restore, and cross-datacenter replication — have high-signal, low-flake coverage; regressions are caught before release, not after.
- Bugs you file consistently include a minimal repro and a suspect location; dev fix cycles get shorter on the surfaces you own.
- AI agents are in daily use inside the framework: writing tests, triaging failures, and evaluated against golden sets so their output is trusted.
- Test-farm flake drops; upgrade/downgrade coverage grows; new features ship with test plans authored before code is complete.
- Engineers across QE, client, and server adopt the patterns, frameworks, and quality bar you introduce.
What this role is not
- This is not manual QA, and it is not a people-manager role unless you later choose that path. Expect to write framework code, attach a debugger to a server, read a compression codec, evaluate an agent run, and leave behind systems other engineers can operate.