Vision CNN SD trainerv005

Browser companion for the XIAO ESP32-S3 FULL VISION ML firmware (v44, or firmware-v005 for config.json class names and debug frames): load the SD card, add images, review and clean them, train the same CNN, then write myWeights.bin back to the card.

Everything runs in this page. Images and weights never leave your computer.

1 Data source

Nothing loaded yet. Pick the root of the SD card (the folder that contains images and header).

2 Classes and data

Class folderImagesTrain / validation

The sketch must be compiled with these values (firmware-v002 and later refuse a weights file of the wrong size). firmware-v003 also reads the class names from header/config.json at boot, but their count must equal NUM_CLASSES, so adding or removing a class means editing NUM_CLASSES and reflashing:

(no classes yet)

Add images from the webcam

● REC

Frames are center-cropped to a square, resized to 240x240 and saved as JPEG straight into images/<class>/ (SD folder) or the zip. Burst takes 10 fresh camera frames back to back, plus the delay you set between them, while the button is red, then saves them. Burst frames are near-duplicates, so move or tilt the object between bursts.

Sample browser

Review steps through one class at a time: mark bad images, then delete them all with one confirmation. Or open a class below and click an image to see its heatmap, delete it, or move it.

3 Train

Model layout

These are compile-time settings in the sketch (INPUT_SIZE, CONV1_FILTERS, CONV2_FILTERS). Copy the lines from section 2 into the sketch after changing them. The 3x3 kernel is fixed in the firmware loops. Images on the card stay 240x240, so changing the size needs no new photos, only a retrain. Cost grows with the square of the input size. Changing the layout discards the model in memory.

Training settings

Idle.

Training never saves by itself. Save the model in section 6 when you are happy with it.

4 Analyze

Confusion matrix

No evaluation yet. It runs automatically when training ends.

Per-class precision and recall

Misclassified images

Click one to inspect it, then delete it or move it to the right class.

5 Infer (live)

Each live frame goes through the same path as a stored image: square crop, 240x240 JPEG, decode, nearest-pixel resize to 64x64.

6 Save

SD folder mode: the existing weights file is copied to myWeights.bin.bak first, then the new one is written. A config.json with the class names and layout is written next to it. firmware-v003 and later read the class names from it at boot (older firmware ignores it). NUM_CLASSES, INPUT_SIZE and the filter counts always come from the compiled sketch.
Zip mode: Save weights keeps the model for the zip. Save .zip downloads header/myWeights.bin (and images if ticked) in the SD card layout.

7 Console

8 Serial monitor

Not connected.

115200 baud, a newline is added to what you send. Close the Arduino IDE serial monitor first. Opening the port can reboot the board, and a reboot may drop the connection: press Connect again. Desktop Chrome or Edge only.

Device view

With the box ticked, firmware-v005 sends its camera JPEG (and a heatmap) every 10th inference, each saved image, and a slow live preview while collecting. Frame lines are shown here instead of in the monitor text. With a model in memory that has the same layout, the page runs the same JPEG through its own copy of the network and compares.

No debug frame yet. Connect, keep the box ticked, then run Infer or a collect mode on the device.

Heatmap: last conv layer, blue = low, red = high. Position is approximate (the last conv map is stretched over the image).

Review images
MARKED BAD

-

Marking deletes nothing. Delete marked asks once, then removes every marked image. Keys: arrows to move, X to mark and advance.