Backends & Guarantee Levels
| Backend | Guarantee | Mechanism |
|---|---|---|
openai() | native | Server-side strict JSON schema |
groq() | native | JSON mode |
deepseek() | native | JSON mode (no schema-strict mode, same tier as groq()) |
fireworks() | native | Server-side JSON schema mode, plus a real token-level GBNF grammar mode |
mistral() | native | Server-side JSON schema mode |
gemini() | native | Server-side JSON schema mode (responseJsonSchema) |
ollama() | constrained | Token-level JSON-schema constraint |
llamaCpp() | constrained | Token-level GBNF grammar (local .gguf via node-llama-cpp) |
anthropic() | best-effort | Prompt + parse + retry |
openRouter() | best-effort | Pass-through to many providers - response_format support varies by underlying model |
together() | native | Server-side JSON schema mode |
cerebras() | native | Server-side strict JSON schema mode |
grok() | native | Server-side JSON schema mode |
openaiCompatible() | best-effort (override to native if your provider enforces it) | Generic factory - see below |
llamaCpp()isconstrainedfor a{ gbnf }input (token-level). For other schema types (Zod / jsonSchema / …) it currently runs a best-effort prompt path until the JSON-Schema→GBNF converter lands - treat those as best-effort despite the nominal level.
fireworks()is the one cloud backend where a{ gbnf }input is not downgraded to best-effort - Fireworks' grammar mode (response_format: { type: "grammar", grammar }) applies the GBNF grammar as a genuine token-level constraint server-side, the same guaranteellamaCpp()gives locally. It reuses theopenaipackage pointed at Fireworks' base URL, so no extra SDK dependency is needed.
openRouter()is deliberatelybest-effort, notnativelike the other cloud backends - it's pass-through across many different underlying providers/models, andresponse_format: { type: "json_schema" }enforcement isn't guaranteed for every model it can route to, only the ones that actually support it themselves.
gemini()uses the official@google/genaiSDK, not an OpenAI-compatible endpoint (unlikefireworks()/mistral()/openRouter()) - Gemini's OpenAI-compat layer is a migration bridge for OpenAI users, not its primary integration path, and doesn't exposeresponseJsonSchema(plain JSON Schema, whattoJsonSchema()already produces) - only the olderresponseSchema(Gemini's own Type-enum OpenAPI-subset shape).
together()/cerebras()/grok()all reuse theopenaipackage pointed at their respective base URLs, same pattern asfireworks()/mistral()/openRouter()- no new SDK dependency. None expose a genuine grammar/constrained-decoding mode, so a{ gbnf }input on any of the three is prompt-only best-effort, same asopenai()/groq().
openaiCompatible()
For any OpenAI-compatible provider shapecraft doesn't name explicitly - baseURL, apiKey, and model are all caller-supplied instead of hardcoded per provider:
import { openaiCompatible } from "@aviasole/shapecraft";
const custom = openaiCompatible({
baseURL: "https://api.some-provider.example/v1",
apiKey: process.env.SOME_PROVIDER_API_KEY,
model: "some-model-id",
guaranteeLevel: "native", // optional - omit to default to "best-effort"
});Defaults to guaranteeLevel: "best-effort" since shapecraft can't verify an arbitrary endpoint actually enforces json_schema server-side. Pass guaranteeLevel: "native" explicitly if you know your provider does, for accurate reporting. Unlike the named backends, there's no environment-variable fallback for the API key - there's no single conventional env-var name for an arbitrary provider, so apiKey is required.
import { openai, groq, fireworks, mistral, gemini, openRouter, deepseek, ollama, anthropic, llamaCpp, together, cerebras, grok, openaiCompatible } from "@aviasole/shapecraft";
const gpt = openai({ model: "gpt-4o-mini" });
const fast = groq({ model: "llama-3.3-70b-versatile" });
const cloudGbnf = fireworks({ model: "accounts/fireworks/models/llama-v3p1-70b-instruct" });
const mist = mistral({ model: "mistral-large-latest" });
const gem = gemini({ model: "gemini-flash-latest" });
const router = openRouter({ model: "openai/gpt-4o-mini" });
const deep = deepseek({ model: "deepseek-v4-flash" });
const local = ollama({ model: "llama3.2" });
const native = llamaCpp({ modelPath: "./models/llama-3.2-3b.gguf" });
const claude = anthropic({ model: "claude-haiku-4-5-20251001", maxRetries: 3 });Model Capabilities
Every built-in backend also exposes capabilities - an explicit, inspectable alternative to duck-typing typeof model.generateStream === "function" for routing logic:
console.log(claude.capabilities);
// { streaming: true, chat: true, structuredOutput: true, toolCalling: true, skillDispatch: true }interface ModelCapabilities {
streaming: boolean; // has generateStream()
chat: boolean; // has chat() - required for turnaround: true
structuredOutput: boolean; // has generate() - always true
toolCalling: boolean; // native provider function-calling, drives generateWithTools() -
// true wherever the backend implements toolCall(); llamaCpp() is
// the one backend without it (local GGUF exposes no tools API)
skillDispatch: boolean; // generateSkillCall()/runSkillLoop() - always true, built on generate()
}capabilities is optional on ShapecraftModel - a custom model implementation that predates this field (or simply doesn't set it) still satisfies the interface unchanged, and model.capabilities is undefined for it. chat?/generateStream? remain the actual methods the core calls; capabilities is just a declared summary of the same information, not a replacement mechanism.
See Tool Calling for toolCalling/generateWithTools(), and Skill-Based Generation for skillDispatch.

