Start free
Grounded in your specs AI Test Case Generator Map the coverage, then write the cases you pick Compare Test Management Pricing Blog Docs Login Start free
AI Test Case Generator

Every tool writes test cases now. Hawzu shows you what they miss.

Point it at a requirement, a pasted spec, or a PDF. Before a single step is written, Hawzu maps up to forty scenarios — grouped by lens, scored for depth per requirement, with the requirements nothing was proposed for named on screen. You pick what's worth writing. It writes those.

Reads RequirementsPasted contextPDFWord (.docx)Markdown & text
Writes Step / expected resultGherkin (BDD)
Hawzu's AI coverage map — scenarios grouped by lens, depth scored per requirement, and one requirement reporting no scenarios proposed

Stage one, on real requirements. Depth per requirement — and R5 reporting nothing proposed, which is the hole a list of finished test cases would never have shown you.

How a run works

Map the coverage. Then write the tests.

Ten finished test cases can't show you a hole — you'd have to already know what should have been there. Forty scenario titles can. So the map is the first thing you see, and it's the thing you approve.

Stage 01

The coverage map

Broad and cheap. One line per scenario, tagged with the lens it came from and the requirement it tests.

Stage 02

The test cases

Deep and narrow. Full steps, written only for the scenarios you ticked — with the whole source re-sent, so the steps still name real fields.

A diagram of the two stages, not a screenshot. The split exists for a plain reason: a finished test case costs roughly ten times what a scenario title costs, so the budget that buys ten cases buys forty lines of coverage — and only one of those is a map.

A generated test case expanded — priority, test type, the requirement it covers, a precondition, and three steps each with its expected result

Stage two, expanded. Priority, test type, the requirement it covers, a precondition, and an expected result on every step — editable before it is saved, and nothing is written until you accept.

What it actually does

Built to be reviewed, not just admired

Machine-written tests are only worth having if a person can check them quickly. Everything here exists to make the check fast.

The map comes before the cases

One call writes about ten finished test cases. The same budget maps forty scenario titles — and forty is enough to see what is missing. You read the map, tick what is worth writing, and only then does it write steps.

It cannot cite a requirement you didn't give it

Each requirement in the run gets an internal label, and the response schema only accepts those labels. A citation to something outside the request isn't discouraged — it is structurally impossible. Work drawn from an attached document cites nothing, rather than guessing.

Ten lenses, not one prompt box

Functional, negative, edge cases and end-to-end — plus six lenses for the things a spec assumes rather than states: state transitions, concurrency, permissions, partial failure, data integrity, and the seams with other systems.

It would rather write less

When the source is too thin, it returns fewer scenarios and says so — instead of padding the list to the number you asked for. A short honest map beats forty plausible titles you then have to disprove one by one.

It knows what you already have

The titles already covering each requirement go into the request, so the map proposes what your suite doesn't have yet. Ask twice and you get new ground the second time, not the same twenty cases again.

Nothing is saved until you say so

Every draft is editable in place — title, priority, preconditions, each step and its expected result. Untick what you don't want. Press Add and the rest land as real, runnable test cases: coded, filed, linked to the requirement they cover, and tagged AI so anyone can see where they came from.

Focus lenses

The half of testing a spec never writes down

A requirement tells you what should happen. It rarely tells you what happens when two people do it at once, or when the payment service times out halfway through.

What the spec states
Functional The behaviour working as described
Negative Invalid input, refused actions, failure states
Edge cases Boundaries, empty states, odd sequences
End to end Journeys crossing the whole behaviour
What the spec assumes
State transitions Every legal move — and the moves that aren't
Concurrency Two actors on the same item at once
Permissions Acting without the role, or on another tenant's data
Failure & recovery Interrupted midway, dependency down, retried
Data integrity Totals matching parts, references surviving deletion
Integration The other side is slow, absent, or returns an error
Measured, not assumed

We ran twenty requirement documents through the generator twice, and had the output judged blind against coverage targets written before any of it existed. Turning the second group of lenses on raised risk coverage from 29.4% to 34.7% — better on eleven documents, worse on one. The gap was never the model's fault: nothing had ever asked it those questions.

Grounded by construction

The rules it can't talk itself out of

Some of these are instructions. The important ones aren't — they're the shape of what the model is allowed to return at all.

Every case traces to your source

The instruction that overrides all others is to invent nothing — no field names, no error strings, no status codes, no data values that aren't in what you supplied.

Gaps are reported, not hidden

Requirements the map proposed nothing for are named on screen, with the two honest explanations: the source doesn't describe them, or they're already covered.

Truncation is disclosed

A source too long to send in full is trimmed from the middle — start and end kept, because acceptance criteria live at the end — and the run tells you exactly how many characters made it.

Your documents are data, never instructions

Attached text and even filenames are treated as material describing a system to test. A file named to look like a command is read as a filename.

It writes test cases. That's all it writes.

It won't create requirements, suites or releases, it won't read your repository, and it won't save a thing you haven't ticked.

The shape of one run

Published ceilings, not discovered ones

You should know where the edges are before you plan around them.

40 scenarios per map 30 by default
20 test cases per run written 10 at a time
10 requirements per run in one brief
3 documents attached one PDF, up to 8 MB
Thursday, 4:40 PM Six requirements, no test cases yet

The spec for Refunds v2 landed this morning as a PDF. You drop it in, tick the six requirements it relates to, leave the risk lenses on, and press Generate. Ninety seconds later you're not reading test cases — you're reading a map:

  • Scenarios mapped 32 across 9 lenses
  • Nothing proposed for REQ-118
  • You tick and write 12 cases, edited, added

The gap on REQ-118 was the real finding. The spec never said what happens to a partial refund on a cancelled order — so nothing could.

Generated cases are ordinary cases

Once you press Add they behave exactly like anything you wrote by hand — they run in executions, count toward requirement coverage, and carry the same codes, folders and history. They keep one small AI tag, so you can always find them again — there's no limbo state, and nothing to approve.

Give it a spec. See what you're missing.

Every feature is free while Hawzu is in early access — no credit card required.

Talk to Us

Tell us about your QA setup. We'll get back to you within 24 hours.

Book a demo

Pick a time that works — we'll confirm by email and send a calendar invite.

Select a date

Available times

Times shown in your timezone: