Beyond chat models, Cobble hosts the pieces you need to make documents searchable and answerable: two OCR models and an embedding model. OCR is billed from your wallet per page; embeddings come out of your plan's monthly embedding allowance, then the wallet. Neither touches your plan's usage budget.
OCR: documents in, text out
Two models, one decision — layout complexity:
| Model | Batch size | Price | Pick it when |
|---|---|---|---|
z-ai/glm-ocr | 200 pages | $0.08/1K pages | Default for clean printed documents |
deepseek/deepseek-ocr-2 | 250 pages | $0.08/1K pages | Mixed layouts, tables, dense pages |
Embeddings: text in, vectors out
| Model | Context | Price | Pick it when |
|---|---|---|---|
google/embeddinggemma-300m | 2K | Plan allowance, then $0.01/1M | Multilingual retrieval and RAG; keep chunks under 2K tokens |
One rule that outranks model choice: embed your documents and your queries with the same model. Vectors from different models live in different spaces; mixing them silently breaks retrieval.
vectors = client.embeddings.create(
model="google/embeddinggemma-300m",
input=chunks, # list of strings
).dataRecipe: PDF → searchable knowledge base
The full pipeline, using only Cobble endpoints plus a vector store of your choice:
# 1. OCR the document (wallet, per page)
text = cobble_ocr("contract.pdf", model="glm/glm-ocr")
# 2. Chunk — start simple: ~800 tokens per chunk, 100 overlap
chunks = chunk(text, size=800, overlap=100)
# 3. Embed the chunks (wallet, $0.01/1M tokens)
vectors = client.embeddings.create(model="google/embeddinggemma-300m", input=chunks)
# 4. Store in your vector DB (pgvector, Qdrant, Chroma...)
db.upsert(zip(chunks, vectors))
# 5. At question time: embed the query, retrieve, generate (plan window)
hits = db.search(embed(question), top_k=8)
answer = client.chat.completions.create(
model="qwen/qwen3.8-27b",
messages=[
{"role": "system", "content": "Answer from the provided context. Cite which excerpt supports each claim."},
{"role": "user", "content": f"Context:\n{format(hits)}\n\nQuestion: {question}"},
],
)Cost intuition for a 1,000-page archive: ~$0.08 to OCR, roughly a cent to embed, and question-answering runs inside your plan window. The expensive part of RAG isn't the pipeline — it's doing it somewhere that meters every step.
Coming soon: a hosted reranker to slot between retrieval and generation, and managed knowledge endpoints that run this whole pipeline for you. The recipe above future-proofs cleanly — the reranker drops in between steps 5's retrieve and generate.
