All toys

Make any PDF readable for AI

Drop a PDF and get clean, structured text back — even from scanned pages. Then split it into bite-sized pieces, ready for a knowledge base.

  1. 1 Upload
  2. 2 Read
  3. 3 Split
0 s processing $0.0000 would be your cost ? free here · – documents left this hour

How it works

Read

AIVAX pulls the text out of your PDF and keeps its structure — headings, lists, tables. Pages that are just images get read with OCR. Fetch & OCR

Split

A model finds where one idea ends and the next begins, so each chunk can answer a question on its own. Your text is never rewritten. Text segmentation

Search

Load the chunks into a collection and your assistant can find and quote the right passage of your document. Collections

For developers API calls from this session, and the code to build it yourself

Calls made by this page

Nothing yet. Read a PDF to see the requests.
API calls
0
Processing units
0
Left this minute
–

Estimates use public list prices before daily allowances. Limits on this toy: 5 documents per minute, 20 per hour. A Turnstile check runs before every call.

Build it yourself

Two API calls: POST /api/v1/web/fetch and POST /api/v1/generations/segment. See the API reference or the Worker behind this page.

# 1. PDF → Markdown. Inline the file as a base64 data URI (10 MB max per item).
curl https://inference.aivax.net/api/v1/web/fetch \
  -H "Authorization: Bearer $AIVAX_API_KEY" \
  -H "Content-Type: application/json" \
  -d "{
    \"contents\": [\"data:application/pdf;base64,$(base64 -w0 manual.pdf)\"],
    \"returnErrors\": true
  }"

# → { "data": { "results": [ {
#       "index": 0,
#       "extractedText": "# Title\n\nParagraph...",   ← Markdown
#       "processingUnits": 7,                      ← what you pay for
#       "jsonProcessingUnits": 0,
#       "error": null } ] } }

# 2. Markdown → segments. `sanitize` skips fragments that are useless for retrieval.
curl https://inference.aivax.net/api/v1/generations/segment \
  -H "Authorization: Bearer $AIVAX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "documents": ["# Title\n\nParagraph..."], "sanitize": true }'

# → { "data": {
#       "result": [ { "index": 0, "count": 8, "segments": ["...", "..."] } ],
#       "usage":  { "processing_units": 2516, "cost": 0 } } }