Learning: Improve Your Form's Reading Accuracy
Learning is a built-in system that helps your form templates get smarter over time. As you scan and correct forms, Nocarta learns from those corrections and automatically improves how it reads your documents. The more you use it, the better it gets.
How to Access Learning
- Open any form template
- Click the Learning button (brain icon) in the settings toolbar
- You'll see the Learning dashboard with overview stats and 7 tabs
Key Concepts
| What We Call It | Technical Name | What It Means |
|---|---|---|
| Practice images | Synthetic data | Artificially generated form images used to train the reading system |
| Difficulty levels | Degradation tiers | How much wear and imperfection is applied to practice images (Easy, Medium, Hard, Adversarial) |
| Knowledge Base | RAG (Retrieval-Augmented Generation) | A collection of patterns and context from your data that helps the system make smarter reading decisions |
| Reading engine | Document reading model | The AI engine that reads and extracts text from scanned images |
| Accuracy profile | Field Accuracy Profile | Per-field confidence thresholds built from your correction history |
| Anchor regions | Alignment anchors | Fixed reference points (logos, headers, borders) used to align scanned images to the template |
| Learning run / Full pipeline | Learning Pipeline | The complete sequence of all learning steps executed in order |
Overview Dashboard
The top of the Learning page shows three key metrics at a glance:
| Metric | What It Means |
|---|---|
| Measurement Coverage | How many of your form instances have been analyzed. Higher is better. |
| Correction Rate | What percentage of scanned fields needed manual corrections. Lower means higher accuracy. |
| Data Readiness | Shows which learning features are available based on how many forms you've processed. Green dots mean ready; amber dots show what's still needed. |
The Seven Tabs
1. Phase 0 — Measurements
This is your starting point. Phase 0 measures the quality of each field in your scanned forms — how clear the text is, how sharp the image is, and how often fields need corrections.
What you can do here:
- View per-field quality scores (contrast, sharpness, how many scans were measured, accuracy)
- Click Measure All Unmeasured to analyze any new form scans that haven't been measured yet
- Adjust sensitivity settings (how strictly blanks are detected, quality classification thresholds)
When to use: After you've scanned several forms and want to see how well each field is being read.
2. Quality (A/B) — Alignment & Image Quality
This tab has two sections:
Alignment & Anchor Regions (Section A): Anchors are fixed reference points on your form (like logos, headers, or borders) that help Nocarta line up scanned images correctly. When scans are properly aligned, each field is read from the right spot on the page.
- View detected anchor regions overlaid on your template image
- Click Auto-Detect Anchors to find anchor regions automatically
- Drag corners to adjust anchor positions manually
- Save your changes with Save Anchors
Image Quality & Accuracy (Section B): Shows how image quality (contrast, sharpness, noise) relates to reading accuracy for each field.
- Color-coded table: green means strong relationship between quality and accuracy, red means problematic
- Preview how image enhancement improves field readability (original vs. enhanced side-by-side)
When to use: When scans are coming in misaligned or certain fields consistently have poor accuracy.
3. Practice Data (E) — Synthetic Image Generation
Generate practice images (synthetic data) with controlled imperfections to train the reading engine. This is especially useful when you don't have many real scanned forms yet. Think of it like giving Nocarta homework to practice reading your form under different conditions.
What you can do here:
- Choose a base image (template background, your best scan, or upload one)
- Pick which fields to include and where their values come from
- Select difficulty levels (degradation tiers):
- Easy: Minor imperfections (slight rotation, light noise)
- Medium: Moderate wear (contrast shifts, slight blur)
- Hard: Heavy wear (creases, stains, shadows)
- Adversarial: Extreme conditions for stress testing
- Click Preview Single to see one example before generating a full batch
- Click Generate Batch to create many training images at once
- Browse generated images in the gallery, filtered by difficulty level
Multiply Real Data: Instead of creating entirely new practice images, you can take your existing real scans and create variations of them with different imperfections. This is often the fastest way to boost accuracy.
- Choose how many variations to create (2x, 3x, 5x, or 10x per original)
- Select which difficulty levels to apply
- Click Multiply Data to generate variations
When to use: When you have fewer than 50 real scans and want to boost training data, or when accuracy on difficult scans (creased, faded, tilted) is low.
4. Accuracy (C) — Confidence & Auto-Correction
Manage how Nocarta uses your correction history to build its accuracy profile (Field Accuracy Profile). This controls when the system is confident enough to fill in values automatically vs. when it flags them for your review.
What you can do here:
- See how your past corrections are being used to improve readings
- Adjust per-field confidence levels (higher = more cautious and asks you to review more, lower = fills in more values automatically)
- Test different confidence settings to preview the impact before applying them
- Click Run Profile Update to refresh the accuracy profile with your latest corrections
When to use: When you want to fine-tune the balance between automatic filling and manual review. Run a profile update after making a batch of corrections.
5. Benchmarks (D) — Compare Reading Engines
Compare different reading engines to find the best one for your specific form type. Requires at least 50 measured instances.
What you can do here:
- View accuracy comparisons across multiple reading engines
- See per-field accuracy breakdowns (which engine reads which field best)
- Compare processing speed and cost
- Click Run Benchmarks to start a new comparison
When to use: After you've accumulated 50+ corrected forms and want to make sure you're using the best reading engine for this form type.
6. Knowledge Base (F/G) — RAG & Contextual Intelligence
The Knowledge Base (powered by RAG — Retrieval-Augmented Generation) collects patterns and context from your form data to help Nocarta make smarter reading decisions. For example, if a field always contains values from a known list (like city names or product codes), the Knowledge Base uses that context to read ambiguous handwriting more accurately.
What you can do here:
- Browse the collected knowledge entries
- See how contextual information is used when reading specific fields
- View evaluation results comparing accuracy with and without the Knowledge Base
- Click Run Refinement to rebuild the Knowledge Base from your latest data
When to use: When certain fields contain specialized values (medical codes, part numbers, legal terms, city names) that benefit from contextual knowledge.
7. Schedule — Automatic Runs
Set up automatic recurring learning runs (scheduled pipeline executions) so your template keeps improving continuously without you having to do anything.
What you can do here:
- Enable/Disable: Toggle automatic runs on or off
- Frequency: Choose how often:
- Daily at 3:00 AM
- Weekly (Sunday at 3:00 AM)
- Monthly (1st of the month at 3:00 AM)
- Custom schedule (cron expression for advanced users)
- Steps: Choose which learning steps to include:
- Field Quality Measurements
- Accuracy Profile Update
- Engine Benchmarks (Reading Model Comparison)
- Knowledge Base Refinement (RAG)
- Data Multiplication (Synthetic Variations)
- If Data Multiplication is enabled, configure how many variations and which difficulty levels (degradation tiers)
- View when the last run happened and when the next one is scheduled
- Click Save Schedule to apply your settings
When to use: Once your form is actively being used and you want continuous improvement without manual effort. A good starting point is weekly runs with Measurements + Profile Update.
Run Full Learning (Pipeline)
The Run Full Pipeline button at the top of the Learning page runs all learning steps in order for this template:
- Measure all unmeasured form scans (Field Quality Measurements)
- Update the accuracy profile from your corrections (Profile Update)
- Compare reading engines (Model Benchmarks — if 50+ forms are available)
- Refresh the Knowledge Base (RAG Refinement)
This is the easiest way to bring everything up to date at once. You'll see live progress updates as each step completes.
Recommended Automation
Follow these steps to get the most out of Learning:
- Start scanning: Upload and process at least 10 forms
- Correct values: Review and fix any reading errors — each correction teaches the system
- Run Phase 0: Click "Measure All Unmeasured" to analyze your data
- Check Quality: Go to the Quality tab and auto-detect anchor regions if alignment is off
- Multiply data: On the Practice Data tab, multiply your real scans to create more training material
- Update profile: Click "Run Profile Update" on the Accuracy tab
- Keep going: As you reach 50+ forms, run benchmarks and Knowledge Base refinement
- Automate: Set up a weekly schedule to keep everything improving on its own
Tips
- More corrections = better accuracy. Every time you fix a reading error, the system learns from it. Don't skip corrections.
- Start with Phase 0. Always measure your data first before trying other features. The overview dashboard shows you what's ready and what needs more data.
- Use data multiplication early. If you only have a few real scans, multiplying them creates enough training material to unlock benchmarks and other advanced features faster.
- Check anchor regions if alignment is off. Misaligned scans are the most common cause of field-level errors. The Quality tab lets you fix this.
- Schedule runs for active forms. Once a form is in daily use, set up weekly or daily automatic runs so accuracy improves continuously without any manual work.
- Progress is shown live. When you trigger any action, you'll see progress updates as the system works. You can navigate away — the work continues in the background.
Frequently Asked Questions
How many forms do I need to start?
You can start with just a few, but the more you have, the better the results. Phase 0 and Quality features work with 10+ forms. Benchmarks require 50+.
How long does a full learning run take?
It depends on the number of forms and fields. For a typical template with 50–100 forms, expect 2–5 minutes. You'll see live progress updates as it works.
Will this change my existing data?
No. Learning only reads your existing forms and corrections to build intelligence that improves future scans. Your existing data is never modified.
What are difficulty levels (degradation tiers)?
When generating practice images (synthetic data), difficulty levels control how much the images are distorted. "Easy" adds minor imperfections (slight tilt, light noise), while "Hard" imitates real-world problems like creases, stains, and shadows. Training with harder levels makes the system better at reading worn or damaged documents.
Can I use Learning on shared templates?
Yes. Any user with edit access to a template can access and run the Learning features. Scheduled runs are set per template and run automatically regardless of who set them up.
What happens if I disable a schedule?
The schedule is paused but not deleted. Your settings are preserved. Toggle it back on and it picks up where it left off. You can always run learning manually at any time using the buttons on each tab.
Need More Help?