Let's give an AI agent its first testing task

Published: · 4 min read

Try a small PDF filename test with Claude Code. Write and run tests, inspect the failure, and stop before changing the code.

Let's give an AI agent its first testing task

Let's give an AI agent its first testing task

You can learn a lot from one small testing task. You need a rule, some code, and a clear stopping point.

Here is a practice task for Claude Code. It checks PDF filenames. You will ask Claude to write tests and run them. Then you will inspect one failure. You will stop before asking for a code fix.

That last step matters. A failing test can be the result you asked for. It shows whether the test found the intended mistake.

What you need

Use a scratch folder that contains only this exercise. You need a terminal and Node.js 24. Check your version with node --version.

You also need Claude Code and a supported account. Follow the official Claude Code quickstart for installation and sign-in. Then check claude --version. Installation and sign-in were not tested in this exercise.

The complete files are below. This example needs no additional packages.

You can also copy the complete public example.

Make the helper

Create a new folder and enter it:

mkdir first-agent-task
cd first-agent-task

Save this as pdf-filename.mjs:

export function isPdfFilename(name) {
  return typeof name === 'string' && name.endsWith('.pdf');
}

The intended rule is simple. Accept filenames ending in .pdf, regardless of letter case. Reject other endings. The helper contains one deliberate mistake: endsWith('.pdf') rejects report.PDF.

An extension is a file ending, such as .pdf. The helper checks the name alone. It does not open a PDF or validate its contents.

Give Claude one bounded task

Start Claude Code inside the scratch folder with claude. Its quickstart shows that command. Paste this prompt into the session:

Read pdf-filename.mjs. The requirement is: accept a filename whose final
extension is .pdf, ignoring letter case. Reject other endings.

Write pdf-filename.test.mjs using Node's built-in node:test runner.
Test invoice.pdf, report.PDF, notes.txt, report.pdf.exe, and an empty name.
Do not change pdf-filename.mjs or install packages.

Run node --test pdf-filename.test.mjs. Report the command, pass and fail counts, and failing case.
Include exact expected and actual values. Then stop.
Do not fix the implementation in this task.

The prompt names the file, rule, test cases, command, and stopping point. It asks for an observable result. Anthropic recommends giving the agent a check it can run. Ask it to show the output.

Read Claude's proposed file change before accepting it. Keep the work inside the scratch folder. If it asks to change the helper, restate the stopping point. After the run, check both filenames. The helper should still match the original. The new test file should hold five cases.

What the test file could look like

Claude may write the tests differently. This complete file gives you a comparison. Save it as pdf-filename.test.mjs if you want to run the example yourself:

import test from 'node:test';
import assert from 'node:assert/strict';
import { isPdfFilename } from './pdf-filename.mjs';

test('accepts a lowercase PDF extension', () => {
  assert.equal(isPdfFilename('invoice.pdf'), true);
});

test('accepts an uppercase PDF extension', () => {
  assert.equal(isPdfFilename('report.PDF'), true);
});

test('rejects a text file', () => {
  assert.equal(isPdfFilename('notes.txt'), false);
});

test('rejects a suffix after .pdf', () => {
  assert.equal(isPdfFilename('report.pdf.exe'), false);
});

test('rejects an empty name', () => {
  assert.equal(isPdfFilename(''), false);
});

Run node --test pdf-filename.test.mjs from that folder. An assertion compares the actual value with the expected value. The runner marks a mismatch as a failure.

In the isolated Node 24.18.0 sample run, four tests passed and one failed. The report.PDF case expected true and received false. Node exited with code 1. The lowercase and rejection cases passed. That is the intended result for this stage.

This example proves the sample test detects the sample mistake. No live Claude Code session was run for this example. Claude's output may differ. Compare its test file and command output with the rule above.

Read the failure before fixing anything

Look for the failing case name and the values false !== true. Check that the failed test points to report.PDF. Then open pdf-filename.mjs and find the case-sensitive ending check.

Also read the passing tests. invoice.pdf should pass. notes.txt, report.pdf.exe, and the empty name should return false. If another case fails, inspect that result before treating the exercise as complete. If all five pass, check whether the helper changed. This task asked for tests only.

Save the command output before starting another task. It records which case failed and what the helper returned. You can compare a later fix against that result.

Now stop. You have a small test that explains a real gap. A separate task can fix the helper and rerun the same tests. Keeping those steps separate makes the first result easier to judge.

Anton Gulin is the AI QA Architect, the first person to claim this title on LinkedIn. He builds AI-powered test automation systems where AI agents and human engineers collaborate on quality. Former Apple SDET (Apple.com / Apple Card pre-release testing). Find him at anton.qa or on LinkedIn.

ai testing · claude code · javascript · software testing

Subscribe

Get notified when I publish something new, and unsubscribe at any time.

Related articles

Read all my blog posts