Grounded in your specs AI Test Case Generator Map the coverage, then write the cases you pick Ask, and check the answer Oracle AI Plain-English questions about your project, answered Compare Test Management Pricing Blog Docs Login Start free
AI in Hawzu

Putting AI in everything is easy. Saying where we didn't is not.

Some answers need judgement. Some have to be the same twice. Every surface here that calls a model says so and names what it read; the rest is arithmetic, printed further down so you can check it. In QA, a number you can't reproduce is worse than no number.

Calls a model
  • Coverage map
  • Existing-coverage judge
  • Test case writing
  • Documentation gaps
  • Readiness narrative
  • Atlas map proposal
  • Rewrite
  • Duplicate defects embeddings
  • Oracle
Arithmetic, and we say so
  • Readiness score
  • Impact tiers
  • Chart picks
  • Observatory insights
  • Flaky detection

same input → same number, every time

Everything on the left names what it read. Everything on the right, you can work out by hand.

The model-backed side

Each one settles a question someone actually asks

Not here are our AI features. Here is the question each one is in the room to answer — and, at the foot of every card, the thing it read in order to answer it.

What should we even be testing here?

It maps the coverage first

Before a single step is written it proposes up to forty one-line scenarios across ten focus lenses, scored for depth per requirement. Ten finished cases can't show you a hole; forty titles can.

reads your specs
More

Do we already test this?

A second pass reads your existing cases

Every proposed scenario is checked against what your repository already covers and ruled covered, partial or not covered — citing the cases it relied on. It errs toward proposing: you can discard a duplicate you can see, not a test you were never offered.

reads your repository

Can it write the ones I picked?

Full cases, only for what you ticked

Steps, expected results, preconditions, priority and the requirement each one covers. It's a preview — nothing is saved until you accept it, and accepted cases land as ordinary test cases with an AI tag.

reads your specs
More

What does our spec fail to say?

It reports the gaps instead of filling them

A behaviour the source names but never gives an outcome for isn't a test case — there's nothing to assert. It comes back as a documentation gap in its own list, deliberately kept out of the drafts so nobody ticks one by accident.

reads your specs

Are we ready to ship?

It writes the readiness narrative

When a release wraps it turns that release's own numbers into the plain-English summary and the top risks, judged against the criteria you wrote down. It puts the scorecard into words — it does not compute it.

reads release metrics your specs
More

What is this product even made of?

It drafts a map of your application

Atlas reads every document in your Canon by heading and proposes the product's areas, screens, routes and journeys — as a draft you review node by node. Nothing reaches the map until a person applies it.

reads your specs
More

Can someone tidy this wording?

Rewrite, tighten, expand or proofread

Pick a mode, see a few variations, preview the exact before-and-after, and apply step by step or all at once. It works on test cases, requirements, releases and defects alike — and one call rewrites one test case, never a batch of them.

reads your repository
embeddings

Has someone already filed this?

It catches the duplicate as you type

Your draft is compared by meaning against every defect in the project, with a match score, so you can close it as a duplicate or jump to a fix someone already found. This one runs on embeddings rather than a language model.

reads your defects
More

Can I just ask?

Oracle answers in whatever shape the answer is

Ask in plain English and it returns the matching records, a chart, an answer from your own documents, or the right form open and ready. It prints every condition it applied, and when nothing can answer you it says so rather than producing something plausible.

reads your project's shape your specs Hawzu's docs
More

All but one call a language model. Duplicate detection compares meaning with embeddings instead — a different technique, named rather than blurred into the rest.

Intake

Everything it reads, you gave it

The chip at the foot of every card above points at one of these. Nothing is grounded in general knowledge about software testing — every model call reads something out of your project, or our own published documentation, and says which.

Reaches a model call

Your specifications

your specs

The PRDs, standards and exit criteria you admitted to Canon. Generation and release calls read them, and cite the document and section they used.

Canon

Your repository

your repository

Your existing test cases, so the second pass can tell you something is already covered — and the exact field you are standing in when you ask for a rewrite. Nothing wider than that.

Test cases

Your project's shape

your project's shape

What your fields are called and the values they hold — folder names, labels, release titles, people. It is what lets Oracle plan a question about your records without reading one.

Oracle

This release's metrics

release metrics

The pass rates, coverage and open blockers of the release in front of it — handed over as data, with an instruction never to invent a number.

Releases

Your defects

your defects

Your project's own native defects, compared by meaning rather than by keyword — so the duplicate surfaces before you file the second one.

Defects

Hawzu's documentation

Hawzu's docs

The in-app assistant answers from our published docs and links the page it used. It reads our documentation, never your content.

Docs

Never general knowledge about software testing, and never another project's data.

In the product

What the boundary looks like on screen

Release 2026.8.1 Completed
Written by a model
Top risks
Arithmetic
72 Needs attention
quality 35
execution 25
defect health 20
requirements 20
One release screen, and both halves of this page inside it. The paragraph was written by a model out of that release's own numbers. The score beside it is the weighted sum, printed further down this page so you can check it.
Boundaries Negative Security +7
REQ-14 Checkout totals 7
REQ-15 Promo stacking 4
REQ-16 Refund window none

Nothing proposed for REQ-16 — the source gives no outcome to assert.

Stage one, before a single step is written. Depth scored per requirement, and the requirement it found nothing for named outright rather than quietly skipped.
New defect
Checkout total wrong when promo applied
Similar defects
DEF-318 Promo code doubles the discount 91%
DEF-402 Basket total ignores promo cap 78%
Surfaced while you're still writing the one that would have duplicated it — by meaning, not by keyword. This is the surface that runs on embeddings.
How do I bulk-import a Postman collection?

“I couldn't find that information in the Hawzu documentation.”

instead of a confident four-paragraph answer
The other half of the same promise, quoted exactly as the assistant says it. Every other answer it gives links the page it came from; this is what happens when there isn't one.
Where we stopped

Five things we deliberately didn't make AI

Every one of these would be easy to ship as an AI feature, and several tools do — the struck-out line on each is the name it would carry. They are arithmetic and rules here, on purpose, because the answer has to be the same twice and you have to be able to argue with it.

“AI-powered release readiness”

The release readiness score

A weighted sum of four dimensions with fixed tier bands. You can recompute it on paper, and two people looking at the same release always get the same number. AI writes the sentence underneath it; it never touches the score.

quality 35 · execution 25 · defect health 20 · requirements 20≥ 80 healthy · ≥ 60 needs attention · ≥ 40 at risk · below, not ready

Those are the bars Hawzu ships with, and a release can carry its own. The bars are yours to set — the arithmetic underneath them is not up for negotiation.

“AI impact analysis”

Atlas impact analysis

A reverse walk over the dependency routes you drew, sorted into three fixed tiers. There is no data to calibrate a weighted risk number against, and a score nobody can predict is worse than a bucket they can.

walk the routes backwards → direct · indirect · unaffected
“AI chart recommendations”

Chart recommendations

A deterministic mapping from what your project looks like to which insights are worth adding. It is the honest version of the phrase every analytics tool prints — the same project always suggests the same charts.

project shape → a fixed list of charts
“AI-generated insights”

Observatory insights

A rule engine over curated thresholds. Every insight states the rule it fired on, so you can disagree with the threshold instead of arguing with a black box.

if metric crosses threshold → say which threshold
“AI flakiness detection”

Flaky test detection

A statistical roll-up over your execution history — a test that passes and fails on unchanged code, counted. No model is involved in deciding what's flaky.

count pass/fail flips per test over its history
The rules it works under

AI you can put in a sign-off

In QA a made-up number is worse than no number. So every surface on the model-backed side works under the same five constraints.

Applies to every model call in Hawzu 05 clauses
  1. 01

    It cites what it used

    A generated case names the document and section it drew on. A docs answer links the page. An uncited claim is visible as an uncited claim.

  2. 02

    It returns fewer rather than padding

    Asked for thirty scenarios from a thin source, it returns the number the source actually supports and says why. The ceiling is a ceiling, not a quota.

  3. 03

    It says when it doesn't know

    No outcome in the spec means a documentation gap, not an invented assertion. No answer in the docs means it tells you so.

  4. 04

    Nothing is written without you

    Generated cases are a preview. Rewrites show the before-and-after. An Atlas proposal is reviewed node by node. There is no background AI writing to your project.

  5. 05

    Your documents are data, never instructions

    Uploaded text and filenames are treated as material describing a system to test. A spec containing “ignore previous instructions” is read as a spec.

Wednesday · 11:20 AM — release review

“Did the AI decide that, or did we?”

Someone points at the amber readiness band and asks the question that should be asked of every AI-shaped number in a review. Here the answer is short: the band came from a weighted sum printed on this page, the paragraph under it was written by a model from those same numbers, and the exit criteria it judged against are a document in your Canon that anyone in the room can open. Three different kinds of answer, and you can tell which is which.

Asked and answered

The questions this page invites

A page that publishes a boundary should be willing to be asked about it. Every answer below is checkable in the product.

What in Hawzu actually uses AI?

Coverage mapping, a second pass that checks proposals against your existing cases, test case writing, documentation gaps, the release readiness narrative, Atlas map proposals, rewrite, duplicate defect detection and Oracle. Duplicate detection is the odd one out — it compares meaning with embeddings rather than a language model, and the page says so rather than blurring it in. Everything else that looks like AI isn't: the readiness score, Atlas impact tiers, chart recommendations, Observatory insights and flaky detection are arithmetic and rules.

Does AI decide whether a release is ready to ship?

No. The readiness score is a weighted sum — quality 35, execution 25, defect health 20, requirements 20 — with tier bands at 80, 60 and 40. Two people looking at the same release always get the same number, and you can recompute it on paper. AI writes the plain-English summary and the top risks underneath it; it never touches the score. A release can carry its own bars, so the thresholds are yours to set — the arithmetic under them is not.

What actually gets sent to a model?

It depends on the surface, and every one of them names what it read on screen. Generation and the readiness narrative send the passages retrieved from your Canon documents, along with that release's own numbers. The coverage judge is shown what your repository already covers. Rewrite sends the field you are editing. Oracle's questions about records send only the shape of your project — what your fields are called and the values they hold — and never the contents of a record. Duplicate detection sends nothing to a language model at all.

Can it change my project on its own?

No. There is no background AI writing to your project. Generated cases are a preview you tick through, and nothing is saved until you accept it. A rewrite shows the exact before-and-after and applies only when you say so. An Atlas proposal is reviewed node by node, and nothing reaches the map until a person applies it.

Can I tell later which test cases came from AI?

Yes, permanently. An accepted case carries an AI badge for the rest of its life, and opening it shows the scenario that produced it and the focus lens that proposed it. A case written by hand carries no marker at all, so a repository that has never used generation shows no extra chrome. It records where the case came from — it is not a review state anyone is waiting on.

Do I need Canon documents for the AI to work?

It depends which surface. Generation grounds itself in your Canon by default and you can turn that off for a one-off, so it works either way. The release readiness narrative is always grounded. Atlas map proposals require Canon outright and refuse rather than guessing at your product's structure without it.

What does it do when it doesn't know?

It says so, in the shape the question arrived in. A behaviour your specification names without giving an outcome comes back as a documentation gap in its own list — deliberately not importable, because there is nothing to assert yet. Asked for thirty scenarios from a thin source, it returns the number the source actually supports and says why. And the documentation assistant answers that it couldn't find the information rather than writing four confident paragraphs.

Are there AI credits or usage limits?

No credits, no balance to top up and nothing to buy — AI is included rather than metered, and you are never shown a token count or a running cost. Each workspace has a generous monthly allowance; nothing appears while you work unless you get close to it, and if a workspace ever does reach it, AI picks up again at the start of the next month. Every feature is free while Hawzu is in early access.

Upstream

The AI is only as good as what it's given

Which is why the surfaces that decide what it reads have pages of their own.

Give it your spec. Check its work.

Every feature is free while Hawzu is in early access — no credit card required.

Talk to Us

Tell us about your QA setup. We'll get back to you within 24 hours.

Book a demo

Pick a time that works — we'll confirm by email and send a calendar invite.

Select a date

Available times

Times shown in your timezone: