Skip to main content
Version: v5

API

Added in: v5.1.0

The models object exposes eight methods. The four that call a model (embed, generate, generateStream and decide) accept an optional model option naming the configured logical model to use; when omitted, the logical name default is used. getDecision() and recordOutcome() read and annotate decisions recorded with persist: true, and, from v5.3.1, calibrate() and getCalibrations() fit and report calibrations from them. None of these four takes the routing options above; getCalibrations() takes an optional { model } filter. Calling a logical name with no configured backend, or asking a backend for a capability it does not support (for example, embeddings from a generation-only backend), throws an error: capability checks run before a backend is called, except the calibrated check on a decision, which a backend can only answer after it returns.

embed()​

models.embed(input: string | string[], options?: EmbedOpts): Promise<Float32Array[]>

Converts one or more strings into embedding vectors. The result is always an array of Float32Array, one per input string, in input order — including when a single string is passed.

import { models } from 'harper';

const [single] = await models.embed('What is Harper?', { inputType: 'query' });
const batch = await models.embed(['first document', 'second document']);
OptionTypeDefaultDescription
modelstring'default'Logical name of a configured embedding model
requiresCapability[]—Capabilities the chosen backend must satisfy; used by routing to select a candidate in the group
inputType'document' | 'query'—Hint for models that distinguish document embeddings from query embeddings (e.g. nomic-embed-text); ignored by models that do not
signalAbortSignal—Cancels the call; composed with the backend's configured requestTimeoutMs

generate()​

models.generate(input: GenerateInput, options?: GenerateOpts): Promise<GenerateResult>

Generates a completion. The input may be:

  • a string — shorthand for a single user message,
  • an array of messages: { role: 'system' | 'user' | 'assistant' | 'tool', content: string },
  • an object { messages, tools?, system? } — the form required to declare tools or pass a system prompt alongside the messages.
const result = await models.generate(
[
{ role: 'system', content: 'You are a terse assistant.' },
{ role: 'user', content: 'What is an HNSW index?' },
],
{ temperature: 0.2, maxTokens: 300 }
);
console.log(result.content);
OptionTypeDefaultDescription
modelstring'default'Logical name of a configured generative model
requiresCapability[]—Capabilities the backend must satisfy (e.g. tools); used by routing. Tools in the input auto-require tools
temperaturenumberbackendSampling temperature, passed through to the backend
maxTokensnumberbackendCompletion token limit, passed through to the backend
responseFormat'text' | 'json' | { schema: object }'text'Structured output. { schema } requests output conforming to a JSON Schema; support varies by backend
toolMode'return' | 'auto''return'How tool calls are handled — see Tool Calling
signalAbortSignal—Cancels the call; composed with the backend's configured requestTimeoutMs

Additional options apply only when toolMode: 'auto'; they are documented in Tool Calling.

GenerateResult​

FieldTypeDescription
contentstringThe generated text
finishReason'stop' | 'length' | 'tool_calls' | 'content_filter'Why generation stopped, normalized across backends
toolCallsToolCall[]Tool calls the model requested, when finishReason is 'tool_calls' (each { id, name, arguments }, with arguments parsed to an object)
usageTokenUsageToken usage reported by the backend (promptTokens, completionTokens, …), when available
traceToolTraceEntry[]Per-tool-invocation trace; only populated by the toolMode: 'auto' loop — see Tool Calling

generateStream()​

models.generateStream(input: GenerateInput, options?: GenerateOpts): AsyncIterable<GenerateChunk>

Identical to generate() but yields the completion incrementally:

let text = '';
for await (const chunk of models.generateStream('Write a haiku about databases.')) {
if (chunk.deltaContent) text += chunk.deltaContent;
}

Each chunk may carry:

FieldTypeDescription
deltaContentstringText appended since the previous chunk
deltaToolCallsPartial<ToolCall>[]Tool-call deltas; a backend may deliver the same tool call across several chunks with partial fields
finishReasonsame values as GenerateResultSet on the final chunk only

Errors detected before the call starts (unknown model name, missing capability) throw synchronously; errors during generation propagate through the iterable.

decide()​

Added in: v5.3.0
models.decide<T>(state: DecideInput, schema: DecisionSchema, options: DecideOpts & { persist: true }): Promise<RecordedDecision<T>>
models.decide<T>(state: DecideInput, schema: DecisionSchema, options?: DecideOpts): Promise<Decision<T>>

Chooses from a closed set of allowed values and returns the chosen value together with a probability distribution over the whole set. For a short, practical introduction, see Start here: typed decisions. Where generate() returns open-ended text, decide() answers a classification, routing, scoring, moderation, or guardrail question with numbers an application can threshold on. It is served by decision backends: a classifier or hosted decision model registered as a custom backend, or the built-in generative adapter, which scores the allowed values from any configured generative model's log-probabilities where the model exposes them and votes over structured completions otherwise.

const decision = await models.decide(ticket.body, {
enum: ['billing', 'refund', 'bug', 'other'],
description: 'Which queue should handle this support ticket?',
});
if (decision.probability < 0.7) await sendToHuman(ticket, decision);

Here decision.value might be 'refund' with a probability of 0.8, and decision.distribution lists every allowed value with its probability, highest first.

state is the input to decide about: a string, or any JSON-serializable object (program state, a record, a message). schema defines the closed set; it is a required argument rather than an option because the schema is what makes the call a decision instead of a generation.

Decision schemas​

A schema is one leaf, or a one-level object of named leaves. Every leaf is a small closed set, so that every backend family — classifiers, cross-encoders, hosted decision models, and language models — can score it:

KindShapeAllowed values
Enum{ enum: [...], description? }2 to 255 distinct values of one type (all strings, all numbers, or all booleans), in the declared order
Boolean{ type: 'boolean', description? }false, true
Bounded integer{ type: 'integer', minimum, maximum, description? }Every integer from minimum to maximum: 2 to 255 values, with safe-integer bounds (32-bit for @decide)
Object{ type: 'object', properties: { [name]: leaf }, description? }One decision per property; at most 32 properties and 500 allowed values across all of them

Any leaf, including an object property, can add noMatch: true to ask for a no-match score as well; the flag is not allowed on an object schema itself.

Deeper nesting, arrays, and free-text extraction are deliberately unsupported: they would split backends into those that can and those that cannot. Object properties are leaves only, and an empty property name or one of __proto__, constructor and prototype is rejected. Descriptions are passed to the backend as task framing; options.instructions adds framing beyond the schema itself.

The set is closed: the distribution is normalized over the allowed values, so an input that matches none of them still produces a confident-looking answer. There are three ways to handle "none of these":

  • Add an explicit value ('other', 'unknown') to a string enum and threshold on probability. It costs nothing extra, but it cannot serve a boolean or integer leaf, and a model may be reluctant to choose the catch-all.
  • Ask a boolean question first ("does this match any of these queues?"), then decide. That is two model calls.
  • Set noMatch: true on the leaf, described next.

No-match scores​

Added in: v5.3.0

A leaf with noMatch: true gets a second number beside its distribution: noMatch, a score from 0 to 1 that the input matches none of the allowed values. The distribution, value and probability keep their meaning. The closest allowed value is still chosen, so the caller decides what to do with a high score, for example send the input to a person instead of acting on value. The two numbers are not a joint distribution: distribution ranks the allowed values as the closest answer whether or not the input matches (the built-in adapter counts every sample's closest value, including samples that said no match), so multiplying it by 1 - noMatch does not give a probability unless the backend documents that guarantee.

const decision = await models.decide(ticket.body, {
enum: ['billing', 'refund', 'bug'],
noMatch: true,
});
if (decision.noMatch > 0.5) await sendToTriage(ticket);
else await route(ticket, decision.value);

Only a backend with the noMatch capability can serve an opted-in schema, so a call routes past backends that cannot produce the score, and fails with a capability error when none can. A backend that returns no score, or one outside 0 to 1, fails the attempt and the next candidate is tried. The score is a backend estimate, not a calibrated probability, unless the backend also claims calibratedNoMatch: an opted-in call reports calibrated: true only when both the distribution and the score are calibrated, and requires: ['calibrated'] on an opted-in schema routes only to such a backend. A schema without the flag is unaffected: it requires nothing new, and no noMatch field appears in its result.

OptionTypeDefaultDescription
modelstring'default'Logical name of a configured decision model
requiresCapability[]—Capabilities the backend must satisfy, e.g. ['calibrated']; used by routing to select a candidate
instructionsstring—Task framing beyond the schema's descriptions, passed to the backend
signalAbortSignal—Cancels the call; composed with the backend's configured requestTimeoutMs
persistbooleanfalsetrue records the decision so its outcome can be reported later; see Recording decisions

Recording decisions​

By default, decide() keeps nothing but its log entry. Like embed() and generate(), each attempt writes one row to the per-call log for observability and usage accounting, and the Decision it returns has no id. That fits most decisions: a route, a moderation verdict, or any threshold an application acts on and moves past.

Pass persist: true when you will want to know later whether the decision was right. The decision is then committed to the hdb_model_decisions system table before the call returns, and the result carries its id. Keep that id with whatever you did: getDecision() reads the decision back, and recordOutcome() records what actually happened.

const decision = await models.decide(ticket.body, { enum: ['billing', 'refund', 'bug', 'other'] }, { persist: true });
await route(ticket, decision.value, decision.id);

Later, once a person has confirmed the right queue, report it against the id route saved on the ticket:

await models.recordOutcome(ticket.decisionId, { truth: { kind: 'value', value: ticket.confirmedQueue } });
Defaultpersist: true
What is storedA row per attempt in hdb_model_calls, buffered and kept for 90 days; a read-only node writes noneThat row, plus a row in hdb_model_decisions that is committed before the call returns and kept for 365 days
What the call returnsDecision<T>, with no idRecordedDecision<T>: a Decision<T> whose id is always present
Recording outcomesNot possiblegetDecision(id) and recordOutcome(id, …)
Cost beyond the model callNoneOne small local commit before the call returns; the row replicates to other nodes in the background
On a read-only nodeWorks as usualRejects with a 503 before any model call, so no provider tokens are spent on a decision that could not be stored
If storing fails after the answerDoes not applyRejects with a 500 without trying another backend, so a storage fault never costs a second model call

Choose when you call. A decision made without persist: true is not stored anywhere an outcome can be attached to, so it cannot be scored later. The @decide directive follows the same rule: it records its decisions only when it names a decision field. A persist that is not a boolean is rejected with a 400 before any model is called.

Decision​

FieldTypeDescription
idstringPresent only when the call passed persist: true: the cluster-unique id of the decision's record in hdb_model_decisions, committed before the decision is returned. Pass it to recordOutcome()
valueTThe chosen value: the most probable outcome for a leaf schema; for an object schema, a map of each property's most probable outcome
probabilitynumberProbability of value (leaf schemas only)
noMatchnumberThe no-match score, for a leaf schema that set noMatch: true; see no-match scores
distribution{ value, probability }[]One entry per allowed value, sorted by descending probability; ties keep the schema's order unless the backend chose one of the tied values, which then leads (leaf schemas only)
fieldsRecord<string, { value, probability, distribution, noMatch? }>Per-property marginals (object schemas only). value is assembled from these marginals and may be a combination no single sample produced
calibratedbooleanWhether the probabilities are calibrated: the backend reports them as calibrated, or a fitted calibration learned from recorded outcomes was applied to every field. The generative adapter's own are not, whether scored from log-probabilities or counted from votes
usageTokenUsageUsage reported by the backend, when available. The generative adapter reports none, because each of its scoring calls or vote samples is recorded as its own scoreChoices or generate call

Harper validates every backend's output against the schema before returning it: value and every distribution entry must be allowed values, the distribution must be complete and sum to one, and value must be a most-probable outcome. A backend that violates this is treated like a failed backend — the attempt is recorded and the next candidate in the fallback group is tried.

A malformed schema, or a state that is not a string or a JSON-serializable object, rejects with a 400 error before any backend is chosen, and writes no analytics row.

getDecision()​

Added in: v5.3.0
models.getDecision<T>(id: string): Promise<DecisionRecord<T> | undefined>

Reads the record of a decision made with persist: true: the schema it was asked over (the input state and instructions are not stored), what was answered, who answered, and whatever has been recorded about it since. The record is committed to hdb_model_decisions before decide() returns, so its id can be looked up right away on the node that made it, after a restart, and on other nodes once replication has delivered it. Returns undefined for an id that does not exist, has expired, or has not reached this node yet. A decision made without persist: true was never stored, so there is nothing to read.

FieldDescription
id, at, expiresAtThe decision's id, when it was made, and when its record and facts expire (365 days after at; a recorded outcome never extends it)
callIdThe hdb_model_calls row of the call that produced it, for correlation; that row is buffered and may be missing after an abrupt shutdown
tenant, appThe tenant and calling resource, when the call carried them
backend, model, signature, configHash, instructionsHashWho answered and under what: the backend, the logical model name, the backend's scoring configuration when it reports one, the identity of the models configuration installed when the call began (a hot reload that lands while a call is in flight is not reflected in that call's record), and the hash of the per-call instructions when any were given
schema, schemaHashThe allowed values the decision was made over (descriptions removed), and the hash of the full schema including descriptions
value, probability, distribution, fields, noMatch, calibratedThe Decision as it was returned
outcomeWhat has been recorded since: { truth?, action?, truthAt?, actionAt? } for a leaf schema, or { fields: { <name>: { … } } } for an object schema

getDecision() and recordOutcome() are administrative, in-process methods with one built-in guard: when the calling request carries a tenant and the record carries a different one, both behave as if the record did not exist. Beyond that they perform no permission check, like every other models method. An application that exposes them to its users must authorize the caller first.

recordOutcome()​

Added in: v5.3.0
models.recordOutcome<T>(id: string, outcome: OutcomeReport): Promise<DecisionRecord<T>>

Records what actually happened for a decision made with persist: true, so later calibration can score predictions against observed truth. A report carries one or two facts, each a tagged state rather than a bare value, so a label that happens to be called 'unknown' is never mistaken for missing information:

FactStates
truth{ kind: 'value', value } (an allowed value of the schema), { kind: 'noMatch' } (the input matched none of them), { kind: 'unknown' }
action{ kind: 'value', value } (the value acted on), { kind: 'noMatch' } (routed as no match), { kind: 'abstained' } (sent to a human by policy), { kind: 'unknown' }
const decision = await models.decide(ticket.body, { enum: ['billing', 'refund', 'bug', 'other'] }, { persist: true });
if (decision.probability >= 0.7) {
await route(ticket, decision.value, decision.id);
await models.recordOutcome(decision.id, { action: { kind: 'value', value: decision.value } });
} else {
await sendToHuman(ticket, decision.id);
await models.recordOutcome(decision.id, { action: { kind: 'abstained' } });
}

When the person's answer is known, record it as the truth, against the id kept with the ticket:

await models.recordOutcome(ticket.decisionId, { truth: { kind: 'value', value: ticket.confirmedQueue } });

For an object schema, report per field: { fields: { queue: { truth: { kind: 'value', value: 'refund' } }, urgent: { action: { kind: 'abstained' } } } }. Each field's facts are recorded and read independently.

Each fact is stored on its own, so recording the truth never touches a previously recorded action, and reports that set different facts never overwrite each other; concurrent reports of the same fact from two nodes converge to one of them. Repeating a report whose state equals what is stored writes nothing. Reporting a different state for the same fact replaces it, so a correction is one more call; { kind: 'unknown' } retracts a fact. Reports are validated against the stored schema: a value must be one of its allowed values, an object schema takes { fields } naming its properties and a leaf schema takes { truth, action }, and a report with no fact is rejected. All of these reject with a 400.

The id must be visible on the node handling the report: an id that does not exist, has expired, or has not replicated to this node yet rejects with a 404. Replication is asynchronous, so an outcome sent to another node immediately after the decision can see that error; record through the node that decided, or retry. Recording an outcome is not a model call: it writes no analytics row and emits no metric. On a read-only node recordOutcome() rejects with a 503 because it is a write, and so does decide() with persist: true, before any model call; a decide() without it stores nothing and works there.

registerBackend()​

Added in: v5.1.15 Changed in: v5.3.0
models.registerBackend(kind: 'embedding' | 'generative' | 'decision', id: string, backend: ModelBackend): void

Registers a custom backend under a logical name, selectable by the model option on later calls. This is the programmatic path for in-process or third-party backends; pair it with models.defineBackend() to build the backend from a few methods. Both are methods on models — reachable as models.registerBackend(...) / scope.models.registerBackend(...) (and likewise for defineBackend), not standalone harper exports. See Custom backends for the full guide.

Errors and timeouts​

  • An unconfigured logical model name throws a not-found error. The error names the missing logical name only — it does not enumerate configured names.
  • A capability mismatch (embedding call to a generation-only backend, tool declarations against a backend without tool support, requires: ['calibrated'] against an uncalibrated decision backend) throws before any request is made. A decision backend that reports a single call as uncalibrated when calibrated was required fails that attempt after the request, and the next candidate is tried.
  • A malformed decision schema or state rejects with a 400 error before any request is made, and is not recorded.
  • On a read-only node, recordOutcome() rejects with a 503 because it is a write, and so does decide() with persist: true, before any request is made; a decide() without it works there. A decide() with persist: true whose record cannot be committed after the backend answered rejects with a 500 without trying another candidate.
  • recordOutcome() rejects with a 404 for an id that does not exist, has expired, or has not replicated to this node yet, and with a 400 for a report that does not fit the stored schema.
  • Each backend supports a requestTimeoutMs configuration field; when set, it is composed with any caller-provided signal so whichever fires first cancels the request.
  • Backend/network failures throw backend-specific errors with sanitized messages.

Every call — successful or failed — is recorded in the model-call analytics.