Agentic Load Testing on K8s: Gap and Build
Agentic load testing teardown: who already sells AI performance testing, the open-source gap they leave, and a Kubernetes reference build with Shamal and k6.
The gap: An open-source, vendor-independent investigation layer that explains load test runs offline, with deterministic saturation detection the model only interprets.
Who buys it: Platform and backend leads at 20-200 person SaaS companies who know they should load test before launches and never have the time.
Status: Repo live · MVP estimate: 6 weeks
The Problem, With Evidence
- A quarter of respondents do not automate load and performance testing, and a third report a greater need for training in it (2024 survey) - German Testing Board survey, via Richard Seidl
- A load-testing vendor says the hard part is translating test intent into a working performance test - Grafana Labs (vendor blog, May 2026)
- Automated traffic surpassed human activity for the first time in a decade, at 51% of web traffic in 2024 - Imperva 2025 Bad Bot Report
- Multimedia bandwidth grew 50% since January 2024, largely from bots scraping images for AI models; 65% of the most expensive traffic comes from bots - Wikimedia Foundation
Market Map: Who Already Does This
| Player | Type | Funding | What it does not do |
|---|---|---|---|
| Grafana k6 / Grafana Cloud k6 | oss + commercial | Acquired by Grafana Labs (2021) | Assistant drafts scripts from natural language or OpenAPI; no published saturation-knee detection or bottleneck explanation |
| Gatling | oss + commercial | Gatling Corp, funding undisclosed | IDE assistant, AI run summary and run comparison are built around Gatling tests and Gatling's own platform; analysis only covers Gatling runs |
| Tricentis NeoLoad | commercial | Tricentis (acquired Neotys, 2021) | Agentic analysis inside an enterprise platform; no verified scenario generation from OpenAPI, not open source |
| BlazeMeter | commercial | Perforce (acquired 2021) | MCP server lets your agent drive tests; the reasoning and knee-finding are left to that agent |
| Artillery | oss + commercial | Y Combinator S21 | Hand-written YAML/JS scenarios; no agentic generation or root-cause step |
| Locust | oss (MIT) | Community project | User behaviour written by hand in Python; no generation or diagnosis |
| Apache JMeter | oss (Apache-2.0) | Apache Software Foundation | GUI/XML test plans built by hand; no native AI generation or diagnosis |
| grafana/mcp-k6 | oss (experimental) | Grafana Labs | Validate and run scripts from an agent; no result analysis or knee detection |
Why Now
Load testing has always had the same problem: the engines are good at generating load and leave everything around it to you. Someone has to write the scenario, someone has to decide what “normal” looks like, and someone has to stare at a latency chart and work out why p95 tripled halfway through the ramp. The German Testing Board’s 2024 survey found a quarter of respondents do not automate load and performance testing at all, and a third want more training in it.
Two things changed in 2025-2026. LLMs got good enough at reading an OpenAPI spec to draft a k6 scenario a human can review in minutes, and tool-calling agents got reliable enough to run a bounded investigation loop: pull summary stats, slice the timeseries, cluster errors, check Prometheus, and write down a ranked hypothesis with evidence. The incumbents noticed. Grafana shipped natural-language k6 script authoring in May 2026, Gatling added an AI run summary in January 2026, and Tricentis launched agentic analysis for NeoLoad in March 2026. Meanwhile the traffic being tested is changing: Imperva reports automated traffic passed human traffic in 2024, and Wikimedia traced a 50% bandwidth jump to AI scrapers.
The Gap
The market map shows where the AI work is going: into each vendor’s own engine and cloud. Grafana’s assistant writes k6 scripts, Gatling’s summarizes Gatling runs, NeoLoad’s analysis lives inside NeoLoad. What nobody offers is the analysis step as an open-source, vendor-independent layer that you can point at your existing k6 results, run fully offline with a local model, and trust because the detection itself is deterministic code.
That is the wedge for agentic load testing as an independent product: keep the industry-standard engine, add an agent that drafts the test and explains the result in a form like “throughput flattened while latency doubled at the knee, errors cluster on connection timeouts, check the pool size before adding nodes”. It does not replace k6 or your thresholds. It replaces the hour of squinting at dashboards that teams skip, and it does it without moving you onto a vendor’s platform.
Reference Architecture
Shamal v0.1.0 is the open-source reference build for this teardown. Shamal v0.1 supports k6 as its only engine, behind an adapter interface meant for Locust or Gatling later. Today it is a CLI that runs anywhere k6 runs: a laptop, a CI runner, or a Pod. The Kubernetes architecture below is how we would run it in a cluster, with the scale-out part provided by Grafana’s open-source k6-operator, not by Shamal itself. Distributed load workers are an explicit non-goal of Shamal v1.
| Component | Role | Runs where |
|---|---|---|
shamal plan | Drafts k6 scenarios from OpenAPI, HAR or an existing k6 script into a fixed template | CI job or laptop |
| Git review | The generated scenario is committed and reviewed like any other code | Your repo |
| k6-operator (Grafana, OSS) | Turns a TestRun resource into parallel k6 runner Jobs for load beyond one machine | Cluster |
| Prometheus | Optional, read-only correlation source for the agent | Cluster |
shamal investigate | Bounded tool-calling loop over the results, ends in a structured finding | CI job or Pod |
shamal report | PR-sized markdown plus a self-contained HTML file | CI job |
The CI gate from the Shamal repo shows the most important design choice: the LLM is never in the pass/fail path. shamal run exits on k6 thresholds alone, and investigation runs afterwards under if: always():
- name: Run the committed scenario
run: shamal run ./perf/checkout.k6.js --results results.json
- name: Investigate and report (runs even when the gate fails)
if: always()
env:
SHAMAL_MODEL: claude-sonnet-5
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
shamal investigate --results results.json
shamal report --results results.json
How We Would Take It to Market
The buyer is not a performance engineer. It is the platform or backend lead at a 20-200 person SaaS company who knows they should load test before a launch and never has the time. They already have k6 or nothing, and they will not adopt a new engine.
That points to three packaging options for a startup built on this pattern:
- Open-core CLI plus a hosted investigation history: the CLI stays free, teams pay for trend storage across runs, comparison between releases, and shared reports. This is the most common devtools model and the easiest to start.
- CI marketplace app: a GitHub or GitLab app that comments on PRs with the investigation, priced per repository or per seat.
- Pre-launch readiness service: a fixed-scope engagement where the tool does the heavy lifting and an engineer signs off. Lower scale, but it pays from day one and produces the case studies the product needs.
The first ten customers are most likely in places where the pain is visible: teams posting “why does my API fall over at N users” in r/devops and r/kubernetes, k6 users who already have scripts but no analysis, and companies preparing a launch or a Black Friday.
What We Learned Building It
A few decisions from building Shamal v0.1.0 that apply to anyone building in this space:
- Keep generated tests as plain engine code. Shamal’s
planoutput is ordinary k6 JavaScript with a provenance header. Delete Shamal and the tests still run. Lock-in is the first objection any team raises about an AI layer over their tests, and this answers it in one sentence. - Constrain the generator. The LLM fills journeys, data, ramp stages and thresholds into a fixed scaffold rather than writing freeform code. That bounds hallucination and guarantees the output parses.
- Make the knee detection deterministic, and let the model only interpret it. Shamal’s saturation check is plain code: take the median p95 of the first 20% of the run as a baseline, then look for the first point where p95 is at least 2x that baseline while virtual users grew 15% or more over the last three samples and requests per second grew 10% or less. It needs at least 8 data points and returns an evidence window, never a cause. The agent’s job is to explain that window, not to find it.
- Budget the agent. The investigation loop is capped at 15 model calls by default and must end with a structured finding: symptom, cited evidence windows, ranked hypotheses with high, medium or low confidence, and next steps. An unbounded agent is a cost line nobody can forecast.
- Air-gap is a feature, not an afterthought. Shamal supports local models through Ollama from day one, makes no telemetry calls, and only connects to the endpoints you configure. Regulated teams that will never send production metrics to a hosted API can still use the whole pipeline.
At v0.1.0 the repo has 122 automated tests passing. What it does not have yet is published benchmark data from real client systems, which is why this teardown shows no performance numbers. We will add them as build logs when we have results we can share.
Unit Economics: What the Reference Build Costs to Run
| Component | Monthly (USD) |
|---|---|
| k6, k6-operator, Shamal (all open source) | 0 |
| LLM calls for investigation (local model via Ollama) | 0 |
| Total | 0 |
Want This Built?
Hire us to build this
We scope, architect and ship the MVP infrastructure on Kubernetes with your team.
Start a buildFrequently Asked Questions
What is agentic load testing?
Agentic load testing uses an AI agent around a normal load engine: it drafts the test scenario from your API description, and after the run it investigates the results with tools (stats, timeseries, error clusters, metrics) and writes a ranked root-cause hypothesis. The engine still generates the load and your thresholds still decide pass or fail.
Does Shamal replace k6?
No. Shamal generates plain k6 JavaScript and runs it through k6. The scenarios run without Shamal installed, and the investigation step also reads k6 results from tests Shamal did not generate. With only a k6 summary export there is no timeseries, so saturation-knee detection needs a run captured through shamal run.
Can I run it on Kubernetes?
Yes, as a CI job or Pod. For load beyond one machine, use Grafana's open-source k6-operator to run k6 as parallel Jobs, then point shamal investigate at the results. Distributed workers are not part of Shamal v0.1.0 itself.
Does it work without sending data to an LLM provider?
Yes. Shamal supports local models through Ollama, has no telemetry, and only connects to the endpoints you configure. The run step needs no LLM at all.
How is this different from Grafana or Gatling AI features?
Those features live inside each vendor's own engine and cloud. Shamal is Apache-2.0, vendor-independent and works offline (v0.1 supports k6; the engine adapter is designed for more), and its saturation detection is deterministic code that the model interprets rather than invents.
Can you build or extend this for my team?
Yes. Use the Start a build link on this page. We can set up the CI gate, the k6-operator path for larger tests, and Prometheus correlation for your stack.