First module of RelGuard — Quality Guard for Relipa

Your test cases stay yours.
Scripts, runs, and bug reports — RelGuard handles those.

RelGuard Test is a VSCode extension that reads exactly the Markdown test cases your team already writes, uses AI to classify which cases can be automated and writes Playwright scripts for those, runs them, captures evidence for every case, and creates Backlog issues on failure. No new tools to learn, no document format to change.

VSCode ≥ 1.85 · Node.js ≥ 18 · Claude Code CLI logged in · Not yet published to Marketplace

100%

Runs inside VSCode — no separate server or cloud needed

1/300

Edit 1 test case out of 300, only that case gets regenerated

Every case

Red-highlighted evidence for both passing and failing cases

demo-customer-app — Visual Studio Code
📥 Import🔄 Sync💡 Suggest➕ Add
IDTest case nameKindSCRIPTROUND 18
TC-001Successful registrationUI✅Pass
TC-002Invalid email addressUI✅Fail
TC-003Age under 18API✅Pass
TC-004Verify with physical deviceManual🖐—

4 test cases · 2 scripts generated · 1 manual

Why RelGuard exists

A registration screen, an Age field, a rule "must be 18 or older". Testers have to read the spec, think of every case, write test cases, then write additional code to run them — and do it all over again at both the API and UI layers.

API: Minimum age = 20

UI: "Age must be 18 or older"

The backend changed its rule; the UI hasn't caught up. Each layer passes in isolation, but users are seeing incorrect information — and there's currently no mechanism to automatically cross-check the two layers.

1

Time-consuming

~30–45 minutes per test case, from thinking of the case to having a working script.

2

Repetitive

The same empty-field / wrong-format / out-of-range patterns, rewritten for every field and every screen.

3

Edge cases get missed

The number of cases depends entirely on what the writer happens to think of that day.

4

No clear picture of what needs re-running

A rule changes from 18 → 20; you have to manually hunt for every test case that involves the Age field.

5

UI and API drift silently

A common bug, hard to catch with manual testing, and usually only visible when a customer spots it.

RelGuard doesn't replace Testers or Developers. RelGuard takes the repetitive parts, so people can focus on the work that requires thought and quality judgment.

Features: everything you need for a complete testing cycle

From importing test cases to landing the issue on Backlog, every step happens inside the same VSCode window.

AI writes scripts — it doesn't invent test cases

Import a Markdown test case file using your team's standard 7-column template. The AI is not allowed to add, remove, or rename any case — it does exactly two things: classify whether each case can be automated or requires manual testing, and write one Playwright script for each automatable case. Manual cases stay in the list with a note explaining why, so your test suite stays complete rather than shrinking to just what a machine can run.

Only regenerate what changed

Every test case has its own hash. Edit TC-015 in a suite of 300 cases and only TC-015 is sent to the AI; the other 299 scripts stay exactly as they are — no rewrites, no wasted tokens. Deleted or reclassified cases have their old scripts moved to a _stale folder — removed from runs but still recoverable.

Live dashboard while running

Run Tests opens a real browser and narrates every Playwright action as it happens (page.goto, locator.click, expect…). Stop / Pause / Resume buttons are available mid-run. If the process crashes or is interrupted, results already collected are saved as a valid round.

running locator.click("Register")
31 pass2 fail13 waiting04:12
⏹ Stop⏸ Pause▶ Resume

Evidence for every case, not just failures

Every test case gets a screenshot with the asserted element highlighted in red and annotated — even passing cases. For API tests, the request and response are rendered as an HTML page with the checked field highlighted. This is submission-ready test evidence, not just debug screenshots.

AI-powered failure analysis

A test fails — click "Analyze failure": the AI returns a root cause, resolution path, action items, and effort estimate, while also distinguishing between a broken test script and a real application bug.

Log bugs directly to Backlog

Creates a Backlog (Nulab) issue with evidence screenshots and a bug report file, description built to your team's exact bug-log template, including occurrence rate like "2/5 rounds (40%)". Re-uploading updates the existing issue — no duplicates.

Suggest missing test cases

AI reviews your existing suite and suggests additions across 4 categories: functional, non-functional, observable performance, and security — each with a reason. You tick which suggestions to keep; they're written back into the original Markdown file.

Test scheduling

Set once / daily / weekly / monthly schedules per function, with automatic end-of-month date handling. If a function already has a run in progress, the next scheduled slot is skipped rather than queued on top.

Parallel script generation

Multiple AI processes run simultaneously; both the number of parallel jobs and cases batched per job are configurable. One job failing doesn't bring down others — completed jobs are written to disk as they finish.

Choose model by task

Three fixed options: Haiku 4.5 (default, lowest cost), Sonnet 5, and Opus 5. Set a workspace default or choose per Sync Test run per function.

Beta · Command Palette

Function Spec lane — generate tests from requirements

Feed in User Requirement + System Requirement documents as Markdown; the AI generates both API tests and UI tests simultaneously for the entire function. Before writing UI scripts, the tool opens the real screen in Chromium to read the DOM and identify the exact widget type for each field (native select, custom combobox, checkbox, radio, date) rather than guessing from descriptions. This lane is temporarily hidden from the sidebar and accessible only via Command Palette — the code remains intact and ready to re-enable.

Four steps, using exactly the test case files your team already has

  1. 01

    Import test cases

    In the RelGuard sidebar, open the function's Test Case panel and click 📥 Import. Select a Markdown file written to the standard 7-column template: Test Case ID · Test Case Name · Severity · Pre-condition · Steps · Data Test · Expected Result. Wrong column headers and the tool reports the exact line number and rejects the import — no silent misreads.

  2. 02

    Sync Test

    Click 🔄 Sync Test, provide the screen URL / API base URL / auth header, and choose an AI model. RelGuard hashes each case, sends only new or changed cases to the AI, and writes out Playwright scripts — cases with IDs starting with API get API tests; the rest get UI tests.

  3. 03

    Run Tests

    One round runs both UI and API tests for that function simultaneously, opening a real browser so you can follow along. Choose the number of workers, run speed, or override the base URL to point at a staging environment without regenerating scripts.

  4. 04

    Review results and log bugs

    Per-round reports include Pass/Fail/Untested counts, timing, and per-case evidence. Aggregate reports across rounds include KPIs, a pass-rate chart, and a Test Execution Summary table. For failed cases, open the detail view, ask the AI to analyze, and file directly to Backlog.

Next time, just edit your test case file and Sync again — RelGuard knows exactly what changed.

Edit one test case — don't pay for the other 299

Per-test-case change detection is the biggest differentiator of RelGuard as a test suite grows.

300 test cases
only TC-015 changed
1 AI call for TC-015
299 other scripts untouched
CriteriaUsual wayWith RelGuard
Data sent to AIRe-reads most of the test case fileOnly the changed portion
AI costScales with file sizeScales with number of changed cases
Unchanged scriptsMay be rewritten by AIPreserved exactly
Review effortReview many changesReview only new parts

This mechanism operates at the per-test-case level for the main flow. In the Function Spec lane, when a rule changes the entire function is reprocessed — because rules displayed across related fields can have cross-cutting effects, and a small diff risks missing related cases.

When a test fails, RelGuard keeps going instead of stopping at "Failed"

Test failure

Automatic evidence

AI analysis

Backlog issue

  • Test failure

    A .md bug report is written at the moment of failure, including error output and evidence image paths.

  • Automatic evidence

    Full-page screenshot with the asserted element highlighted in red and an annotation label.

  • AI analysis

    Root cause · Resolution path · Action items · Effort estimate — plus a clear statement of whether this is a broken test script or a real application bug.

  • Backlog issue

    Issue type Bug, summary [TC-ID] case name, description built to your team's exact bug-log template, with evidence image and bug report file attached.

# Bug Report (tester, reporter)
## Test Environment
## Preconditions
## Steps
## Actual results
## Expected results
## Occurrence rate     → 2/5 rounds (40%) — failed in rounds 3, 5
## Evidence

# Automated analysis (RelGuard AI)
## Root cause / Resolution / Action items / Effort estimate

# Bug fix (Developer)
## Root cause / Changes made / Scope of impact / Merge Request / Self Review

Preview the full content that will be uploaded, right inside VSCode, before clicking send.

Expected impact

The numbers below are targets RelGuard aims for, based on a survey of the current testing workflow.

CriteriaCurrentWith RelGuardExpected
Writing tests + scripts~60–90 min / 100 cases~20–30 min / 100 cases~75% reduction
Review on changeReview many sectionsFocus on changed parts~80% reduction
Missed edge casesHeavily experience-dependentAI-assisted suggestions~50% reduction
Tests executed / day~50 cases~150 cases~200% increase

These are estimates from the Hackathon proposal. Actual figures will be measured during the Q1/2027 pilot.

Roadmap

RelGuard follows the order Test → Process → Security. Each guard only begins after the previous one has been piloted on a real project with clear evaluation results.

  1. Q4/2026In progress

    Polish for pilot readiness

    Migrate all secrets to SecretStorage · set up CI running typecheck + unit tests on multiple platforms · solve the scheduler dependency on VSCode being open · automate .vsix builds on tag · choose an internal distribution channel · select 1–2 pilot projects.

  2. Q1/2027Planned

    Pilot and measure real impact

    Run on 1–2 real projects, capture bugs that only surface in production-like environments, re-measure all figures in the table above with real data, then make a go/no-go decision for wider rollout.

  3. Q2/2027Planned

    Expand + launch RelGuard Process

    Roll out to more projects if the pilot succeeds. Begin designing and building the RelGuard Process MVP — quality gates for Requirements and Code Review, reusing the existing Core Engine, Report Renderer, and sidebar.

  4. Q3/2027Direction

    Launch RelGuard Security

    Define specific scope (SAST, dependency scan, DAST…), build MVP and pilot. This milestone is currently directional — it requires a UC Spec and technical scope before implementation can begin.

RelGuard — Quality Guard for Relipa
   ├── RelGuard Test      → Existing test cases · Function Spec
   ├── RelGuard Process   → Quality gate for Requirements / Code Review
   └── RelGuard Security  → Security vulnerability scanning

An open question at management level: whether real customer spec and test case data is permitted to pass through the Claude API during the pilot phase. Without a decision here, the Q1/2027 milestone may be pushed back.

Install

About 5 minutes if your machine already has Node and Claude Code CLI.

RelGuard Test has not been published to the VSCode Marketplace. Install using the .vsix file downloaded directly here — VSCode may show a warning about an unsigned extension, which is normal for internal extensions.

VSCode ≥ 1.85Node.js ≥ 18Claude Code CLI logged inPlaywright Chromium
  1. 1Download the .vsix file using the button below.
  2. 2In VSCode, open Extensions (Ctrl+Shift+X) → ⋯ menu in the top corner → Install from VSIX…
  3. 3Select the downloaded file; VSCode will prompt to Reload when done — click Reload.
  4. 4A shield icon appears in the activity bar. Click it; the sidebar has 5 sections: Project · Functions · Testing · Schedule · Settings.

First steps after installing

  • Run claude interactively once inside the project directory to accept the trust dialog — RelGuard calls Claude via CLI so this step is required.
  • Sidebar → Functions → ➕ create your first function (one screen or one feature = one function).
  • Sidebar → Testing → Test Case → 📥 Import your Markdown test case file, then 🔄 Sync Test.
Download RelGuard Test 0.1.0 (.vsix)

Size ~4.4 MB · Compatible with VSCode ≥ 1.85

Frequently asked questions

No. RelGuard calls Claude through the Claude Code CLI already on your machine, reusing that existing session. No API key is hardcoded or stored in the extension.

Yes — prompts sent to Claude include test cases, requirement documents, DOM information, and bug report content, so that data passes through the Claude API. This is why using RelGuard on real customer data is pending a data-governance decision at management level. The Backlog API Key is stored using vscode.SecretStorage, encrypted by the operating system.

Not in the main flow. The AI is not allowed to add, remove, or rename any case in your imported file — it only classifies (automatable / manual) and writes scripts. If you want AI suggestions for new cases, there is a dedicated "Suggest missing test cases" button, and every suggestion requires you to tick it before it is saved.

They stay in the test case list, labeled manual with a note explaining why (physical device interaction, visual inspection, simulated network loss…). Your test suite stays complete — not trimmed down to only what a machine can run.

No. The schedule timer lives inside the extension process and only runs while VSCode is open — it is not an OS cron job, and there is currently no alert when a scheduled slot is missed. Splitting this out to run independently is on the Q4/2026 roadmap.

Yes. The only limitation on Windows is the Pause button mid-run (requires a POSIX signal) — the Stop button works normally.

Not currently. RelGuard has shifted its design: input is now either a Markdown test case file using the team's standard template, or Markdown requirement documents in the Function Spec lane. Supporting additional formats remains out of scope.

Chromium (Desktop Chrome) via Playwright. Firefox, WebKit, and mobile devices are not yet supported, nor is responsive testing or visual comparison.

Ready to let AI write the scripts for you?

Download the extension, import the test case file you already have, click Sync Test once, and see the result.

VSCode ≥ 1.85 · Node.js ≥ 18 · Claude Code CLI · Playwright Chromium