Skip to main content
Version: v5

Schema

Harper uses GraphQL Schema Definition Language (SDL) to declaratively define table structure. Schema definitions are loaded from .graphql files in a component directory and control table creation, attribute types, indexing, and relationships.

Overview​

Added in: v4.2.0

Schemas are defined using standard GraphQL type definitions with Harper-specific directives. A schema definition:

  • Ensures required tables exist when a component is deployed
  • Enforces attribute types and required constraints
  • Controls which attributes are indexed
  • Defines relationships between tables
  • Configures computed properties, expiration, and audit behavior

Schemas are flexible by default — records may include additional properties beyond those declared in the schema. Use the @sealed directive to prevent this.

A minimal example:

type Dog @table {
id: Long @primaryKey
name: String
breed: String
age: Int
}

type Breed @table {
id: Long @primaryKey
name: String @indexed
}

Loading Schemas​

In a component's config.yaml, specify the schema file with the graphqlSchema plugin:

graphqlSchema:
files: 'schema.graphql'

Keep in mind that both plugins and applications can specify schemas.

Type Directives​

Type directives apply to the entire table type definition.

@table​

Marks a GraphQL type as a Harper database table. The type name becomes the table name by default.

type MyTable @table {
id: Long @primaryKey
}

Optional arguments:

ArgumentTypeDefaultDescription
tableStringtype nameOverride the table name
databaseString"data"Database to place the table in
expirationInt—Seconds until a record goes stale (useful for caching tables)
evictionInt0Additional seconds after expiration before a record is physically removed
scanIntervalInt(expiration + eviction) / 4Seconds between eviction scans
replicateBooleantrueEnable replication of this table
auditBooleanlogging.auditLog for ordinary tablesEnable the table's transaction log; new tables with @fullText require explicit true
cacheControlString—Cache-Control header value emitted on anonymous GET/HEAD 200/304 responses for this table
randomAccessFieldsBooleanstorage.randomAccessFieldsPin this table's record encoding

expiration, eviction, and scanInterval

These three arguments work together to control the full lifecycle of a cached record:

  • expiration — When elapsed, a record is considered stale. The next request for a stale record triggers a fetch from the source. The record may still be served while revalidation is in progress.
  • eviction — Additional time after expiration before the record is physically removed from the table. Setting eviction > 0 lets you serve the stale record while revalidation happens and controls how long after expiration the data is kept on disk.
  • scanInterval — How often Harper scans the table for records to evict. Defaults to one quarter of expiration + eviction.

You can provide a single expiration value and all three behaviors share the same TTL. To tune them independently:

# Expire after 5 minutes, evict after 1 hour, scan every 10 minutes
type WeatherCache @table(expiration: 300, eviction: 3300, scanInterval: 600) {
id: ID @primaryKey
temperature: Float
}

How scanInterval Determines the Eviction Cycle​

scanInterval determines fixed clock-aligned times when eviction runs. Harper divides the clock into evenly spaced anchors based on the interval, calculated in the server's local timezone. As a result:

  • The server's startup time does not affect when eviction runs.
  • Eviction timings are deterministic and timezone-aware.
  • For any given configuration, the eviction schedule is the same across restarts and across servers in the same local timezone.

Example: 1-hour expiration — default scanInterval = 15 minutes (one quarter of expiration). Eviction schedule:

00:00, 00:15, 00:30, 00:45, 01:00, ...

If the server starts at 12:05, the first eviction runs at 12:15 — not 12:20. The schedule is clock-aligned, not startup-aligned.

Example: 1-day expiration — default scanInterval = 6 hours. Eviction schedule:

00:00, 06:00, 12:00, 18:00, ...

Eviction with Indexing​

Eviction removes non-indexed record data, but it does not remove a record from its secondary indexes. If an evicted record matches a search query, Harper fetches the full record from the source on demand to satisfy the query. This means indexes remain fully functional even when most of the data has been evicted.

randomAccessFields​

Added in: v5.1.0

Encodes this table's records as typed random-access structures, so a single field can be read without decoding the whole record. It suits tables whose records carry a stable set of fields with stable types, and should be left off for wide, sparse, or variably typed schemas. See Storage Tuning — Record Encoding for the trade-off and for how to check where a table sits against the structure bound.

Declaring the argument pins this table's encoding when the table is created: the table keeps it regardless of the storage.randomAccessFields setting, which otherwise applies to every table that does not declare it. That cuts both ways — a fleet-wide change to the global setting will not move a table that pins its own encoding, and editing the argument later does not repin an existing table. Omit the argument when creating a new table to follow the global setting instead.

type Reading @table(randomAccessFields: true) {
id: ID @primaryKey
sensorId: String @indexed
celsius: Float
recordedAt: Int
}

Examples:

# Override table name
type Product @table(table: "products") {
id: Long @primaryKey
}

# Place in a specific database
type Order @table(database: "commerce") {
id: Long @primaryKey
}

# Auto-expire records after 1 hour (e.g., a session cache)
type Session @table(expiration: 3600) {
id: Long @primaryKey
userId: String
}

# Disable replication for this table explicitly
type LocalRecord @table(replicate: false) {
id: Long @primaryKey
value: String
}

# Combine multiple arguments
type Event @table(database: "analytics", expiration: 86400) {
id: Long @primaryKey
name: String @indexed
}

Database naming: Since all tables default to the data database, when designing plugins or applications, consider using unique database names to avoid table naming collisions.

Replication: Replication is enabled by default for all tables. Note that if you disable replication on a table and re-enable it later, it will not catch-up on previous writes during when the replication was disabled.

cacheControl​

Added in: v5.2.0

The cacheControl argument sets a Cache-Control header value emitted on anonymous (unauthenticated) GET/HEAD 200/304 responses for this table. It is designed for tables whose content is public and safe for shared caches (CDNs, reverse proxies) to store.

type Product @table(cacheControl: "public, max-age=60") @export {
id: Long @primaryKey
name: String
price: Float
}

Key semantics:

  • Anonymous reads only. The header is emitted only when the request carries no authenticated principal. Authenticated responses instead receive an identity floor of Cache-Control: private, no-cache (plus Vary: Authorization/Cookie) — the table declaration is not inherited by authenticated reads.
  • Explicit opt-in required. Anonymous readability alone does not cause Harper to emit shared-cache headers, because a table gated on request attributes (IP, headers, etc.) could leak across a URL-keyed cache. The cacheControl declaration is the explicit statement that the content is safe to cache publicly.
  • Never on 401 responses. A rejected request always gets Cache-Control: private, no-cache regardless of any declaration.
  • Visible in describe_table. The value is stored with the table schema and surfaces in the Operations API's describe_table output.

The declaration can equivalently be made as a static cacheControl property on an exported JavaScript resource class:

export class Product extends tables.Product {
static cacheControl = 'public, max-age=60';
}

See REST Headers / Cache-Control for the full caching behavior and how Cache-Control interacts with authentication headers.

@export​

Exposes the table as an externally accessible resource endpoint, available via REST, MQTT, and other interfaces.

type MyTable @table @export(name: "my-table") {
id: Long @primaryKey
}

The optional name parameter specifies the URL path segment (e.g., /my-table/). Without name, the type name is used.

@export alone does not serve HTTP traffic — REST must also be enabled for the application, either explicitly with rest: true in its config.yaml or by Harper's built-in default for a component directory that has no configuration file at all. See REST Overview / Tables and Their Automatic Endpoints for the endpoints an exported table produces and the exact conditions.

@export is a routing directive, not an access control

Omitting @export removes the REST/MQTT route for a table (callers get 404), but it does not protect the data. The table still exists in the database and remains accessible through the Operations API and SQL, subject to RBAC, to administrators and roles with the required operation and table permissions. For table-level confidentiality, omit the table from a role's grants or set its table-level read permission to false rather than relying on the absence of an export route. To protect individual fields, configure attribute_permissions as a whitelist: list every field the role may read and omit the restricted field (or list it with read: false). Also account for the filter side-channel on exported Resources.

@sealed​

Prevents records from including any properties beyond those explicitly declared in the type. By default, Harper allows records to have additional properties.

type StrictRecord @table @sealed {
id: Long @primaryKey
name: String
}

@hidden (Type Directive)​

Suppresses the type from introspectable surfaces — MCP tool descriptors and the OpenAPI document. The table still exists; data is still queryable through Harper's other interfaces subject to RBAC. @hidden is a metadata-visibility directive, not an access-control mechanism: use table-level role permissions to control access to the table and attribute_permissions to control access to individual fields.

type InternalConfig @table @hidden {
id: Long @primaryKey
value: String
}
@hidden does not restrict data access

@hidden only suppresses a type or field from generated API specs and MCP tool schemas. The underlying data remains available on every surface through which the table is reachable — REST, MQTT, and GraphQL when the table is exported, plus SQL and the Operations API — subject to the user's role permissions. Do not use @hidden as a confidentiality control. To restrict access to a table, omit it from a role's grants or set its table-level read permission to false. To restrict access to an individual field, configure role attribute_permissions as a whitelist: list every permitted field and omit the restricted field (or list it with read: false). Also account for the filter side-channel on exported Resources.

@hidden is also available as a field directive to suppress individual attributes.

Documenting Types and Fields​

Harper picks up GraphQL's standard triple-quoted docstrings on type and field definitions. Docstrings flow through to:

  • MCP — Table.description (consumed as a prefix on every verb-tool description) and inputSchema.properties[*].description on derived tool schemas
  • OpenAPI — components.schemas[*].description, per-property description, and the path-level description for every verb on the resource
"""
Product catalog row — what shows up in the storefront listing,
search, and inventory feeds. One row per SKU.
"""
type Product @table @export {
"""
Stock keeping unit — globally unique across catalogs.
"""
sku: String! @primaryKey

"""
Display name shown in the storefront.
"""
name: String!

"""
Retail price in cents (USD).
"""
priceCents: Int!
}

Docstrings on @hidden fields are dropped from the descriptive surfaces alongside the field itself.

Trust model. Docstrings reach LLMs and public OpenAPI consumers verbatim. Treat them as code: don't put secrets, internal-only commentary, or speculative prose in them. Use @hidden to suppress fields that shouldn't surface publicly.

Field Directives​

Field directives apply to individual attributes in a type definition.

@primaryKey​

Designates the attribute as the table's primary key. Primary keys must be unique; inserts with a duplicate primary key are rejected.

type Product @table {
id: Long @primaryKey
name: String
}

If no primary key is provided on insert, Harper auto-generates one:

  • UUID string — when type is String or ID
  • Auto-incrementing integer — when type is Int, Long, or Any
Changed in: v4.4.0

Auto-incrementing integer primary keys were added. Previously only UUID generation was supported for ID and String types.

Using Long or Any is recommended for auto-generated numeric keys. Int is limited to 32-bit and may be insufficient for large tables.

@indexed​

Creates a secondary index on the attribute for fast querying. Required for filtering by this attribute in REST queries, SQL, or NoSQL operations.

type Product @table {
id: Long @primaryKey
category: String @indexed
price: Float @indexed
}

If the field value is an array, each element in the array is individually indexed, enabling queries by any individual value.

Null values are indexed by default (added in v4.3.0), enabling queries like GET /Product/?category=null.

@embed​

Added in: v5.1.0

Automatically computes an embedding vector for the attribute whenever the source field is written, using a configured embedding model:

type Document @table {
id: Long @primaryKey
text: String
embedding: [Float] @embed(source: "text", model: "default")
}
  • source: the name of the field to embed. Must be a declared field on the same type, passed as a string literal.
  • model: the logical name of a configured embedding model, passed as a string literal.

The attribute type must be [Float]. The attribute is automatically indexed with an HNSW vector index, so it is immediately searchable by similarity; an explicit @indexed on the same attribute is allowed only if it is also HNSW.

Write semantics:

  • Creating a record with the source field, or updating the source field, computes the vector before the write commits (with inputType: 'document'). A failure to compute the embedding fails the write.
  • An update that does not touch the source field leaves the vector unchanged.
  • Setting the source field to null sets the vector to null.
  • Replicated writes and audit-log replays do not re-embed — the vector travels with the record, and only the node that accepted the original write calls the model.
  • Changed in: v5.3.0 On a read-only node the embedding is not computed: writes cannot commit there, and a caching table serves a record filled from its source without computing the vector or caching the row.

Multiple @embed attributes on one type are computed concurrently.

@fullText​

Added in: v5.3.0 Engine: RocksDB

Creates a BM25-ranked full-text derived index from one or more stored text fields. Declare the directive on a nullable FullText field; that field's name identifies the index.

type Product @table(database: "catalog", audit: true) {
id: ID @primaryKey
name: String
description: String
tags: [String]
catalogSearch: FullText @fullText(fields: [{ name: "name", weight: 3 }, { name: "description" }, { name: "tags" }])
}

The FullText field is query-only. It is not stored with records or exposed as a selectable record property. Add another FullText field to create another independent index. Full-text indexes require RocksDB and an audited table. They are local derived state rebuilt from committed table data rather than authoritative record storage. See Full-Text Search Configuration for the complete field contract, filter metadata, phrase and prefix storage, synonyms, highlighting, and rebuild behavior.

@decide​

Added in: v5.3.0

Automatically makes a typed decision about the source field whenever a write carries it, using a configured decision model (Start here: typed decisions walks through a first one), and stores the chosen value on the attribute and, optionally, its probability on a second attribute:

type Ticket @table {
id: Long @primaryKey
body: String
route: String @decide(source: "body", values: ["billing", "refund", "bug", "other"], confidence: "routeConfidence")
routeConfidence: Float @indexed
urgent: Boolean @decide(source: "body")
severity: Int @decide(source: "body", minimum: 1, maximum: 5)
}
  • source: the name of the field to decide about. Must be a declared field on the same type, passed as a string literal. A string is passed to the model as text; an object is passed as program state.
  • values: the allowed strings, for a String attribute: 2 to 255 distinct values, passed as a list of string literals.
  • minimum, maximum: the inclusive range, for an Int attribute: at most 255 values, passed as integer literals.
  • model: the logical name of a configured decision model, passed as a string literal. Defaults to "default".
  • confidence: the name of a nullable Float field on the same type that receives the probability of the chosen value. Optional.
  • decision: the name of a nullable String field on the same type that receives the id of the decision's durable record, so an outcome can later be recorded for the stored value. Optional; see Recording outcomes for a decided attribute.
  • instructions: task framing sent to the model with every decision, passed as a string literal. Optional.

The attribute type selects the decision schema: String with values is an enum, Boolean takes no further arguments, and Int with minimum and maximum is a bounded integer. Other attribute types are rejected. The attribute and the confidence field must be nullable, because a null source clears them, and neither can be the primary key or @computed. The closed set is validated when the schema loads, so a bad values list fails deployment rather than the first write. The attribute is not indexed implicitly: add @indexed to query by the value, and index the confidence field to query by probability, as above.

Write semantics match @embed:

  • Creating a record with the source field, or updating the source field, calls models.decide before the write commits and stores the value and the probability from that one decision. A failure to decide fails the write, so the two are never written separately. The hook sees the write payload before table validation, so a write that is later rejected has already paid for its decisions.
  • An update that does not touch the source field leaves both unchanged. The decided attribute and its confidence field are ordinary attributes: a write that carries them without the source stores them as given, so a stored probability records a model call rather than proving one. The decision field is not: see Recording outcomes for a decided attribute. Three writes change the source without deciding again, as with @embed: a tracked-instance edit that assigns the source after update() and saves, a CRDT operation payload on the source, and a write sent with x-replicate-from: none. All three keep the previous pair, and the previous decision id, beside the new source, including when the new source is null.
  • Setting the source field to null sets both to null, and the decision field too when the directive names one.
  • Replicated writes and audit-log replays do not decide again — the value and probability travel with the record, and only the node that accepted the original write calls the model. Upgrade every node that accepts writes before adding @decide to a schema: a node that does not know the directive commits the source without the pair, and the other nodes store what it sent.
  • Changing values, the model or the instructions applies to writes from then on. Existing rows keep their values, and no index is rebuilt.
  • On a caching table (one populated from a sourcedFrom source), the decision runs when a record is filled from its source, after the reader's GET has already returned the source record. A failure there aborts only the cache write: the reader sees no error, the row is not cached, and the next read fetches from the source and decides again. The first read of an uncached row does not wait for the decision, so do not rely on it carrying the decided value or its decision id. On a node running in read-only mode no write can commit, so neither @decide nor @embed calls a model there, and a caching table serves the source record without storing it.

Every @decide attribute on a type is decided concurrently, alongside any @embed attributes. When one of them fails, the others are signalled to stop and the write fails once they have settled. A derived field has exactly one writer: two directives cannot name the same attribute, confidence or decision field, and a directive cannot use another directive's output as its source.

The probability is whatever the configured backend reports, or, from v5.3.1, its calibrated form when a correction applies. With the built-in generative adapter under its default scoring: auto, an attribute the generative model can score (OpenAI scores up to 20 allowed values from token log-probabilities) gets a normalized log-probability score from one scoring call; an attribute it cannot score, a model that exposes no log-probabilities, and a call that every scoring candidate declines are voted over samples completions instead; a decline after the scoring request was sent is billed before the vote, while a decline decided before any request, such as more allowed values than the backend can score, costs nothing. Without calibration neither number is a measured probability, so a threshold such as routeConfidence < 0.7 is a review-queue rule rather than a guarantee. Cost follows the same split: the Ticket type above makes three decisions on every write that carries body, three scoring calls when every attribute scores or fifteen completions at the default samples when they vote, including a write that table validation later rejects. On a caching table the count is per fill rather than per write: every fill of an uncached or expired record decides again, even when the source has not changed, and a fill whose decision fails is not cached, so the next read pays again. Any reader of an uncached key spends a decision, so limit who can read a caching table that decides. A vote also needs a generative candidate with the structuredOutput capability unless the decision entry sets requireStructuredOutput: false; a group with none fails the write before any completion is requested. For a high write rate, route the directive's model name to a scoring-capable generative model or a single-call decision backend, or lower samples.

A component can replace the default decider for an attribute with the table's static setDecideAttribute(name, decider), for example tables.Ticket.setDecideAttribute('route', decider). The decider receives the write payload and returns { value, probability }, or null to clear the attribute and its confidence and decision fields; when the directive names a decision field the decider may also return the id of a models.decide call it made with persist: true, and a missing id stores null, meaning no recorded decision. An override for a directive without a decision field has no use for persist: true: nothing could reach the record. The value must be one the directive allows, the probability is required when the directive names a confidence field, and the override survives a schema reload. Its second argument carries a signal that aborts when another hook of the same write fails, so a decider that forwards it to its own model call lets the write fail promptly.

Recording outcomes for a decided attribute​

A directive records its decision only when it names a decision field, the same choice persist: true makes for a direct models.decide call. Without one, nothing is stored beyond the model call's entry in hdb_model_calls. With one, the directive decides with persist: true and the field holds the decision's id, which models.getDecision() reads and models.recordOutcome() records what actually happened for:

type Ticket @table {
id: Long @primaryKey
body: String
route: String @decide(source: "body", values: ["billing", "refund", "bug", "other"], decision: "routeDecision")
routeDecision: String
}
import { models, tables } from 'harper';

async function recordRoute(ticketId, route) {
const ticket = await tables.Ticket.get(ticketId);
if (!ticket?.routeDecision) return undefined;
return models.recordOutcome(ticket.routeDecision, { truth: { kind: 'value', value: route } });
}
  • The decision field is written by the directive only. A write that carries it is rejected with a 400 unless the same write carries the source, whose new decision then replaces it; a replicated write or an audit-log replay stores the id it carries. An x-replicate-from: none request is not exempt. The check sees the write payload only: code that assigns the field on a tracked instance after update() and saves is not checked, like the tracked-instance source edit above, so trusted code should leave the field to the directive.
  • The id is provenance, not authorization. recordOutcome checks only the request's tenant, and a request with no tenant can report on any decision, so the code that records an outcome must authorize the caller for the record and its tenant itself. Code in a resource runs in a trusted context and skips table permissions, so reading the record there authorizes nothing: check that the caller may report on this ticket before calling a function like recordRoute.
  • The field holds the latest decision only. A later write that carries the source replaces it, and a null source or a delete leaves the previous decision unreferenced; capture the id with the action you take if the outcome arrives later.
  • Association is best effort, not atomic: the decision is committed before the record, so a write that fails after its decision (validation, a conflict, an abort) leaves an unreferenced decision row. Unreferenced rows expire with the 365-day retention, and a record can outlive its decision, in which case getDecision returns nothing and recordOutcome returns 404.
  • Decisions replicate separately from records, so on another node a freshly replicated record can name a decision that has not arrived yet; retry on 404, or record the outcome on the node that made the decision.
  • Existing records gain a decision value only when a later write decides again. Deploy the new version to every node that accepts writes before adding decision: to a schema: an older node rejects the argument.

@createdTime​

Automatically assigns a creation timestamp (Unix epoch milliseconds) to the attribute when a record is created.

type Event @table {
id: Long @primaryKey
createdAt: Long @createdTime
}

@updatedTime​

Automatically assigns a timestamp (Unix epoch milliseconds) each time the record is updated.

type Event @table {
id: Long @primaryKey
updatedAt: Long @updatedTime
}

@expiresAt​

Marks a field as the record's absolute expiration time (Unix epoch milliseconds). Once the timestamp is in the past, the record is treated as expired: it is hidden from reads and physically removed by the eviction sweep. Useful for session, token, or per-record cache-entry tables where each record carries its own expiry.

type Session @table {
id: ID @primaryKey
token: String
expiresAt: Long @expiresAt
}

The @expiresAt field is authoritative over the table-level expiration default, in both directions — a per-record value may extend a record's expiration past (since 5.2), or expire it before, the table default. When a record omits the field, the table default (if any) applies.

  • The field must be an absolute timestamp (Unix epoch milliseconds), not a duration. Negative values are ignored and fall back to the table default.
  • A full-record put that omits the field clears it (the record then follows the table default, or never expires if there is none). A patch of other fields preserves it — so a session-renewal write must re-include expiresAt.
  • On caching tables (with a source), set the per-record TTL via context.expiresAt in the source get() rather than an @expiresAt field.

@hidden (Field Directive)​

Suppresses the field from MCP tool descriptors and the OpenAPI document. The attribute still exists in the table and can be returned through every surface on which the table is reachable (REST GET, MQTT, and GraphQL when exported, plus SQL and the Operations API), subject to the user's field permissions. Use this for fields that should not appear in generated specs or tool schemas, not to restrict data access.

type Customer @table {
id: Long @primaryKey
name: String

"""
Internal — do not surface to external consumers.
"""
creditScore: Int @hidden
}

@hidden is a metadata-visibility directive, not access control: attribute_permissions on roles remains the field-level data-access enforcement mechanism. To prevent a role from reading a field value, configure attribute_permissions as a whitelist: list every permitted field and omit the restricted field (or list it with read: false). Also account for the filter side-channel on exported Resources.

Relationships​

Added in: v4.3.0

The @relationship directive defines how one table relates to another through a foreign key. Relationships enable join queries and allow related records to be selected as nested properties in query results.

@relationship(from: attribute) — many-to-one or many-to-many​

The foreign key is in this table, referencing the primary key of the target table.

type RealityShow @table @export {
id: Long @primaryKey
networkId: Long @indexed # foreign key
network: Network @relationship(from: networkId) # many-to-one
title: String @indexed
}

type Network @table @export {
id: Long @primaryKey
name: String @indexed # e.g. "Bravo", "Peacock", "Netflix"
}

Query shows by network name:

GET /RealityShow?network.name=Bravo

If the foreign key is an array, this establishes a many-to-many relationship (e.g., a show with multiple streaming homes):

type RealityShow @table @export {
id: Long @primaryKey
networkIds: [Long] @indexed
networks: [Network] @relationship(from: networkIds)
}

@relationship(to: attribute) — one-to-many or many-to-many​

The foreign key is in the target table, referencing the primary key of this table. The result type must be an array.

type Network @table @export {
id: Long @primaryKey
name: String @indexed # e.g. "Bravo", "Peacock", "Netflix"
shows: [RealityShow] @relationship(to: networkId) # one-to-many
# shows like "Real Housewives of Atlanta", "The Traitors", "Vanderpump Rules"
}

@relationship(from: attribute, to: attribute) — foreign key to foreign key​

Both from and to can be specified together to define a relationship where neither side uses the primary key — a foreign key to foreign key join. As with the to-only form above, the result type must be an array: Harper resolves it by searching the target table's to attribute for matches, using this record's from attribute (instead of its primary key) as the search value.

type OrderItem @table @export {
id: Long @primaryKey
orderId: Long @indexed
productSku: Long @indexed
products: [Product] @relationship(from: productSku, to: sku) # matches products by sku, not primary key
}

type Product @table @export {
id: Long @primaryKey
sku: Long @indexed
name: String
}

Schemas can also define self-referential relationships, enabling parent-child hierarchies within a single table.

Computed Properties​

Added in: v4.4.0

The @computed directive marks a field as derived from other fields at query time. Computed properties are not stored in the database but are evaluated when the field is accessed.

type Product @table {
id: Long @primaryKey
price: Float
taxRate: Float
totalPrice: Float @computed(from: "price + (price * taxRate)")
}

The from argument is a JavaScript expression that can reference other record fields.

Computed properties can also be defined in JavaScript for complex logic:

type Product @table {
id: Long @primaryKey
totalPrice: Float @computed
}
tables.Product.setComputedAttribute('totalPrice', (record) => {
return record.price + record.price * record.taxRate;
});

Computed properties are not included in query results by default — use select to include them explicitly.

Computed properties that read other tables use the same trusted server-side authorization context as other direct tables or databases calls. Do not expose protected cross-read data through a computed property when access depends on the caller.

Computed Indexes​

Computed properties can be indexed with @indexed, enabling custom lookup strategies such as composite keys. Use @fullText for native BM25-ranked text search and an HNSW [Float] field for vector indexing:

type Product @table {
id: Long @primaryKey
tags: String
tagsSeparated: String[] @computed(from: "tags.split(/\\s*,\\s*/)") @indexed
}

When using a JavaScript function for an indexed computed property, use the version argument to ensure re-indexing when the function changes:

type Product @table {
id: Long @primaryKey
totalPrice: Float @computed(version: 1) @indexed
}

Increment version whenever the computation function changes. Failing to do so can result in an inconsistent index.

Vector Indexing​

Added in: v4.6.0

Use @indexed(type: "HNSW") to create a vector index using the Hierarchical Navigable Small World algorithm, designed for fast approximate nearest-neighbor search on high-dimensional vectors.

type Document @table {
id: Long @primaryKey
textEmbeddings: [Float] @indexed(type: "HNSW")
}

Embedding vectors can also be computed automatically at write time from a text field with the @embed directive, which creates the HNSW index implicitly.

Query by nearest neighbors using the sort parameter:

let results = Document.search({
sort: { attribute: 'textEmbeddings', target: searchVector },
limit: 5,
});

HNSW can be combined with filter conditions:

let results = Document.search({
conditions: [{ attribute: 'price', comparator: 'lt', value: 50 }],
sort: { attribute: 'textEmbeddings', target: searchVector },
limit: 5,
});

Changed in: v5.2.0 — Conditions combined with a vector sort are evaluated during graph traversal (predicate-aware search): the search keeps exploring until it has enough matching nearest neighbors, instead of finding the nearest candidates first and then dropping the ones that fail the filter. With a selective filter this is the difference between a full result set and an under-filled one. When a companion condition is very selective, Harper instead computes exact distances over just the records matching that condition, which is both exact and faster than traversing the graph.

Filtered Vector Search with a Function Predicate​

Added in: v5.2.0

A vectorFilter function on the query participates in the traversal the same way, for predicates that are not expressible as conditions:

let results = Document.search(
{
sort: { attribute: 'textEmbeddings', target: searchVector },
vectorFilter: (record) => record.tenantId === context.user.tenantId && record.status === 'published',
limit: 10,
},
context
);

vectorFilter is available from the JavaScript API only (it cannot be expressed in a REST query string). The function receives the candidate record and must return a boolean — true to include the record in results, false to exclude it (it still routes traversal either way). It must be synchronous, side-effect free, and fast — it can run once per candidate record visited during traversal (verdicts are memoized per query). Records passed to it are frozen.

Row-Level Access Control with Explicit Filters​

Added in: v5.2.0

Use a rowFilter function on a search target or subscription request when access depends on each record. Attach the filter in an operation override after deciding that the request itself is allowed to proceed. This keeps authorization ahead of query planning and table scans while still filtering the returned or delivered rows:

function canReadReport(record, context) {
const user = context.user;
if (user?.role?.permission?.super_user) return true;
return user?.username != null && record.ownerId != null && record.ownerId === user.username;
}

export class Reports extends tables.Reports {
get(target) {
const context = this.getContext();
if (target.isCollection) {
target.rowFilter = canReadReport;
} else if (!canReadReport(this, context)) {
return new Response(null, { status: 404 });
}
return super.get(target);
}

search(target) {
target.rowFilter = canReadReport;
return super.search(target);
}

subscribe(request) {
request.rowFilter = canReadReport;
return super.subscribe(request);
}
}

rowFilter is available only from the JavaScript API; clients cannot set it through REST or QUERY request data. It receives the candidate record and the live request or subscription context. It must be synchronous, side-effect free, and fast. Ordinary object and array records are passed as shallow read-only views. Throwing an error or returning a promise aborts a query; on a subscription it terminates the stream. The filter is enforced across normal index and range searches, OR queries, HNSW traversal, source-revalidated records, subscription snapshots, and live row events.

rowFilter does not apply to a direct primary-key get. The example performs that check in get() after the default instance-loading flow has loaded the record. It therefore assumes the default loadAsInstance behavior. A false-mode handler must explicitly load any record data it needs before making the same decision.

For vector queries, rowFilter participates in HNSW traversal. A caller therefore receives the k nearest matching records rather than "nearest k, minus filtered records." rowFilter can apply to every candidate visited; prefer indexed query conditions for predicates that can be expressed declaratively.

rowFilter applies only when a non-raw put or invalidate event carries an authoritative row value. Delete tombstones, value-less invalidations, raw events, and published messages may not contain the complete current record. A subscription that needs such events can also provide an eventFilter:

subscribe(request) {
request.rowFilter = canReadReport;
request.eventFilter = (event, context) => {
const rowEvent = event.type === 'put' || event.type === 'invalidate';
if (!request.rawEvents && rowEvent && event.value != null) return true;
const username = context.user?.username;
const ownerPrefix = username == null ? null : `${encodeURIComponent(username)}:`;
return (
ownerPrefix != null &&
(rowEvent || event.type === 'delete') &&
String(event.id).startsWith(ownerPrefix)
);
};
return super.subscribe(request);
}

This example assumes report IDs are created with the same encodeURIComponent(username) + ':' owner prefix, allowing value-less events to be authorized without loading the deleted or invalidated record. Encoding the owner segment prevents one username from being a raw string prefix of another owner's IDs.

eventFilter receives a shallow read-only view of every non-control event and the live context. It has the same synchronous and fail-closed requirements as rowFilter and is composed with it for authoritative row events. When a subscription has a rowFilter, non-row events are withheld unless an eventFilter explicitly accepts them. Transaction and reload control events continue to pass so the subscription remains coherent.

Live subscriptions are periodically re-authorized with a freshly loaded user. This recheck reruns the operation-level allowRead grant; it does not call a custom subscribe() override again. The rowFilter and eventFilter callbacks receive the refreshed context for subsequent events. If an application-specific connection grant must terminate an existing stream when revoked, keep that revocable grant in an allowRead override composed with super.allowRead; admission logic that runs only in subscribe() cannot provide that teardown behavior.

The legacy allowRead, allowUpdate, allowCreate, and allowDelete hooks are deprecated operation-level gates. When permission checking is active, Harper's standard instance flow evaluates the relevant hook once for the operation rather than once per row. Built-in table handlers with loadAsInstance = false likewise use one request/collection-scoped verdict; custom false-mode handlers remain responsible for authorization unless they delegate to those built-in handlers.

Put application authorization and access-control decisions that need the complete target and context in operation overrides such as get, put, and delete. Collection operation overrides run before query planning or table scans and can attach a rowFilter or indexed conditions when an admitted request needs row-level narrowing. In the default single-record get flow, the record instance is loaded before the operation override runs, so its fields are available through this.

See the RequestTarget and SubscriptionRequest references for the filter properties.

Tuning Filtered Traversal​

Filtered traversal is bounded by a visit budget of ef * filterExpansion nodes (filterExpansion defaults to 24). If the budget is exhausted before the result list fills — which happens when the filter matches only a tiny fraction of records — the search returns the matches found so far rather than erroring. Both knobs can be set per query:

let results = Document.search(
{
sort: { attribute: 'textEmbeddings', target: searchVector, ef: 200, filterExpansion: 40 },
vectorFilter: (record) => record.category === 'rare',
limit: 10,
},
context
);

Raise filterExpansion (or ef) to trade latency for recall under selective function predicates. Condition-based filters rarely need tuning: very selective conditions are automatically diverted to the exact-scan strategy instead of graph traversal.

Filtering by Distance Threshold​

To return only records whose distance to a target vector is below a threshold, place target directly on the condition (alongside comparator and value). This returns matches within the threshold without using sort:

let results = Document.search({
conditions: {
attribute: 'textEmbeddings',
comparator: 'lt',
value: 0.1,
target: searchVector,
},
});

This form is useful when you want to bound result quality by a similarity cutoff rather than ranking by similarity.

Selecting the Distance​

Use the special $distance field in select to include the computed distance from the target vector in returned records:

let results = Document.search({
select: ['name', '$distance'],
sort: { attribute: 'textEmbeddings', target: searchVector },
limit: 5,
});

$distance is available in both sort-based ranking and conditions-based threshold queries.

Per-Query Search Options​

The sort descriptor (and threshold condition) accepts options that tune an individual query:

let results = Document.search({
sort: { attribute: 'textEmbeddings', target: searchVector, distance: 'dotProduct', ef: 200 },
limit: 5,
});
  • distance — overrides the index's distance function for this query: "cosine", "euclidean", or "dotProduct" (dotProduct Added in: v5.1.0).
  • ef Added in: v5.1.0 — overrides the search exploration budget for this query. Higher values improve recall at the cost of latency.

Changed in: v5.1.0 — When a query passes no ef and the index does not explicitly configure efConstructionSearch (or efConstruction), the search budget auto-scales with the size of the index, so recall holds as the table grows instead of decaying with a fixed budget.

HNSW Parameters​

ParameterDefaultDescription
distance"cosine"Distance function: "cosine" (negative cosine similarity), "euclidean", or "dotProduct" (added in v5.1.0)
efConstruction100Max nodes explored during index construction. Higher = better recall, lower = better performance
M16Preferred connections per graph layer. Higher = more space, better recall for high-dimensional data
optimizeRouting0.5Heuristic aggressiveness for omitting redundant connections (0 = off, 1 = most aggressive)
mLcomputed from MNormalization factor for level generation
efConstructionSearchauto-scaledMax nodes explored during search. When unset, auto-scales with index size (see above); setting it (or efConstruction, which seeds it) fixes the budget
quantization—"int8" stores vectors quantized to int8 (added in v5.1.0, see below)
filterExpansion24Visit-budget multiplier for filtered (predicate-aware) search: a filtered query visits at most ef * filterExpansion nodes (added in v5.2.0, see above)

Example with custom parameters:

type Document @table {
id: Long @primaryKey
textEmbeddings: [Float] @indexed(type: "HNSW", distance: "euclidean", optimizeRouting: 0, efConstructionSearch: 100)
}

Note: this parameter was previously documented as efSearchConstruction; the option name Harper reads is efConstructionSearch.

Changed in: v5.1.0 — Changing efConstructionSearch on an existing index no longer triggers a rebuild; it only affects searches. Structural parameters (distance, M, efConstruction, quantization) still rebuild the index when changed.

Vector Quantization​

Added in: v5.1.0

quantization: "int8" stores the index's vectors quantized to 8-bit integers, substantially reducing index size and memory traffic:

type Document @table {
id: Long @primaryKey
textEmbeddings: [Float] @indexed(type: "HNSW", quantization: "int8")
}

Graph navigation runs on the quantized (approximate) distances. For nearest-neighbor sort queries, Harper re-ranks the results against the full-precision vectors stored on the records, restoring exact ordering and exact $distance values. Distance-threshold (lt/le) queries currently filter on the approximate distance.

Field Types​

Harper supports the following field types:

TypeDescription
StringUnicode text, UTF-8 encoded
Int32-bit signed integer (−2,147,483,648 to 2,147,483,647)
Long54-bit signed integer (−9,007,199,254,740,992 to 9,007,199,254,740,992)
Float64-bit double precision floating point
BigIntInteger up to ~300 digits. Note: distinct JavaScript type; handle appropriately in custom code
Booleantrue or false
IDString; indicates a non-human-readable identifier
AnyAny primitive, object, or array
DateJavaScript Date object
BytesBinary data as Buffer or Uint8Array
BlobBinary large object; designed for streaming content >20KB
FullTextQuery-only full-text index declaration; valid only with @fullText and never stored

Added FullText in v5.3.0

Added BigInt in v4.3.0

Added Blob in v4.5.0

Large integers: Long vs BigInt​

Despite the name, Long is bounded by the JavaScript safe-integer range — values must satisfy |value| < 2^53 (9,007,199,254,740,991, i.e. Number.MAX_SAFE_INTEGER). A Long attribute rejects integers beyond that range, so it is not a full 64-bit type.

For true 64-bit integers — IDs, counters, or timestamps that can exceed 2^53 — use BigInt, and send the value as an actual bigint (for example, via CBOR or MessagePack) rather than a JSON number. A JSON number above 2^53 has already lost precision before Harper receives it, so the larger value cannot be recovered. BigInt attributes, including @indexed range queries, store and order distinct values above 2^53 correctly.

Arrays of a type are expressed with [Type] syntax (e.g., [Float] for a vector).

Blob Type​

Added in: v4.5.0

Blob fields are designed for large binary content. Harper's Blob type implements the Web API Blob interface, so all standard Blob methods (.text(), .arrayBuffer(), .stream(), .slice()) are available. Unlike Bytes, blobs are stored separately from the record, support streaming, and do not need to be held entirely in memory. Use Blob for content typically larger than 20KB (images, video, audio, large HTML, etc.).

See Blob usage details below.

Blob Usage​

Declare a blob field:

type MyTable @table {
id: Any! @primaryKey
data: Blob
}

Create and store a blob using createBlob():

let blob = createBlob(largeBuffer);
await MyTable.put({ id: 'my-record', data: blob });

Retrieve blob data using standard Web API Blob methods:

let record = await MyTable.get('my-record');
let buffer = await record.data.bytes(); // ArrayBuffer
let text = await record.data.text(); // string
let stream = record.data.stream(); // ReadableStream

Blobs support asynchronous streaming, meaning a record can reference a blob before it is fully written to storage. Use saveBeforeCommit: true to wait for full write before committing:

let blob = createBlob(stream, { saveBeforeCommit: true });
await MyTable.put({ id: 'my-record', data: blob });

Any string or buffer assigned to a Blob field in a put, patch, or publish is automatically coerced to a Blob.

When returning a blob via REST, register an error handler to handle interrupted streams:

export class MyEndpoint extends MyTable {
static async get(target) {
const record = await super.get(target);
let blob = record?.data;
if (!blob) return record;
blob.on('error', () => {
MyTable.invalidate(target);
});
return { status: 200, headers: {}, body: blob };
}
}

Dynamic Schema Behavior​

When a table is created through the Operations API or Studio without a schema definition, it follows dynamic schema behavior:

  • Attributes are reflexively created as data is ingested
  • All top-level attributes are automatically indexed
  • Records automatically get __createdtime__ and __updatedtime__ audit attributes

Dynamic schema tables are additive — new attributes are added as new data arrives. Existing records will have null for any newly added attributes.

Use create_attribute and drop_attribute operations to manually manage attributes on dynamic schema tables. See the Operations API for details.

OpenAPI Specification​

Tables exported with @export are described via an /openapi endpoint on the main HTTP server associated with the REST service (default port 9926).

GET http://localhost:9926/openapi

This provides an OpenAPI 3.x description of all exported resource endpoints. The endpoint is a starting guide and may not cover every edge case.

Renaming Tables​

Harper does not support renaming tables. Changing a type name in a schema definition creates a new, empty table — the original table and its data are unaffected.

  • JavaScript API — tables, databases, transaction(), and createBlob() globals for working with schema-defined tables in code
  • Data Loader — Seed tables with initial data alongside schema deployment
  • REST Querying — Querying tables via HTTP using schema-defined attributes and relationships
  • Resources — Extending table behavior with custom application logic
  • Storage Algorithm — How Harper indexes and stores schema-defined data
  • Configuration — Component configuration for schemas