Skip to main content
Version: v5

Querying Full-Text Indexes

Added in: v5.3.0 Engine: RocksDB

Use the query-only FullText field name as a condition attribute in Table.search().

const products = Product.search({
conditions: [{ attribute: 'catalogSearch', comparator: 'matches', value: 'waterproof trail shoes' }],
limit: 20,
});

Table.search() returns an async iterable. Consume it with for await...of, or return it directly from a custom Resource to stream the results.

Results are ordered by BM25 relevance. The FullText field, not one of its source fields, identifies the search target.

Match modes​

ComparatorBehaviorRequired index option
matchesMatches any analyzed query term.—
matches_allRequires every analyzed query term.—
matches_phraseMatches analyzed terms in order as a phrase.positions: true
matches_prefixMatches records whose analyzed terms start with the final query term.surfaceTerms: true
matches_fuzzyMatches terms within one edit, including a transposition.—
matches_fuzzy_prefixCombines one-edit fuzzy matching with a final-term prefix.surfaceTerms: true

Every comparator also has a negated form: not_matches, not_matches_all, not_matches_phrase, not_matches_prefix, not_matches_fuzzy, and not_matches_fuzzy_prefix.

An empty string is invalid and returns 400. Whitespace-only text or text reduced to no terms by the analyzer, such as a query containing only enabled stop words, returns an exact empty result.

A query containing a negated full-text condition must also contain a non-negated full-text condition on the same index. A structured record condition alone does not satisfy this requirement. The positive full-text condition bounds the candidate set; the negated condition filters it.

const products = Product.search({
operator: 'and',
conditions: [
{ attribute: 'catalogSearch', comparator: 'matches', value: 'waterproof' },
{ attribute: 'catalogSearch', comparator: 'not_matches', value: 'leather' },
],
});

Restricting source fields​

By default, a condition searches every source field in the index. Use fields to restrict it:

const products = Product.search({
conditions: [
{
attribute: 'catalogSearch',
comparator: 'matches_phrase',
value: 'trail running',
fields: ['name', 'description'],
},
],
});

Every requested field must belong to the index and be readable by the caller. Without fields, the caller must be allowed to read every source field in the index. Harper rejects an external REST or Operations API query with 403 if any searched source field is not readable. $highlights never bypasses this check.

Direct calls through tables or databases run in a trusted server-side context. From an authenticated custom resource, pass the context received by the static resource method to Table.search() and set checkPermission: true in the search options to enforce the caller's table and source-field permissions. Context propagation alone does not request the table permission check.

Score and ordering​

Select $score to include the BM25 score:

const products = Product.search({
conditions: [{ attribute: 'catalogSearch', comparator: 'matches', value: 'trail shoes' }],
select: ['id', 'name', '$score'],
limit: 20,
});

Full-text results use descending relevance order. The only explicit full-text sort is $score descending; other sort keys and reverse iteration are rejected because they would discard the native ranking.

An exact count requested through Table.search() does not report a full-text total after authorization, structured filtering, and current-record validation. It returns recordCount: null and recordCountExact: false. REST behaves the same when its interface enables exactCount; otherwise Prefer: count=exact is downgraded to an estimated count.

Highlights​

Highlighting returns source-field fragments and matching character spans. The index must configure highlighting and mark at least one source with highlight: true. A query can also search sources without that option; those sources can affect matching and ranking but are omitted from $highlights.

Selecting $highlights turns highlighting on for the query:

const products = Product.search({
conditions: [{ attribute: 'catalogSearch', comparator: 'matches_phrase', value: 'trail running' }],
select: ['id', 'name', '$score', '$highlights'],
});

You can also set includeHighlights: true on the condition. A result has this shape:

{
"id": "shoe-1",
"name": "Waterproof trail running shoe",
"$score": 2.41,
"$highlights": {
"name": [
{
"valueIndex": 0,
"spans": [{ "start": 11, "end": 24 }],
"fragments": [
{
"text": "Waterproof trail running shoe",
"start": 0,
"spans": [{ "start": 11, "end": 24 }]
}
]
}
]
}
}

Offsets are UTF-16 half-open ranges. Value-level spans address the original source value. A fragment's start addresses that value, while spans inside the fragment address its returned text. Harper returns offsets rather than HTML so the caller controls rendering and escaping.

Combining conditions​

Full-text conditions can be combined with structured filters using and:

const products = Product.search({
operator: 'and',
conditions: [
{ attribute: 'catalogSearch', comparator: 'matches', value: 'trail shoe' },
{ attribute: 'category', comparator: 'equals', value: 'footwear' },
{ attribute: 'available', comparator: 'equals', value: true },
],
limit: 20,
});

When category and available are declared in the index's filterFields, Harper includes these conditions in the Tantivy query. Tantivy removes non-matching IDs before Harper loads records. The metadata is score-neutral, and Harper rechecks the original conditions against current records before returning them.

The same query remains valid when an attribute is not declared in filterFields; Harper applies that condition after full-text search. Small, complete result sets from Harper secondary indexes can also be passed to Tantivy automatically as candidate IDs. This path does not require filterFields.

An or group may combine conditions from the same full-text index:

const products = Product.search({
operator: 'or',
conditions: [
{ attribute: 'catalogSearch', comparator: 'matches_phrase', value: 'trail running' },
{ attribute: 'catalogSearch', comparator: 'matches', value: 'hiking boot' },
],
});

An or group cannot mix full-text and ordinary record conditions, and one query cannot combine different full-text indexes. Full-text conditions must name a field directly on the queried table; relationship and nested-property paths are not supported. Run separate queries when any of these boundaries is required.

Freshness controls​

A full-text index is derived from committed table changes. Each condition accepts:

OptionDefaultDescription
maxIndexLagMilliseconds3000Maximum accepted upper bound on index lag. Set to 0 to require current coverage.
waitForIndexMilliseconds0How long to wait for acceptable coverage, from 0 through 30000.

For read-after-write behavior, require current coverage and allow a bounded wait:

const products = [];
for await (const product of Product.search({
conditions: [
{
attribute: 'catalogSearch',
comparator: 'matches',
value: 'new product',
maxIndexLagMilliseconds: 0,
waitForIndexMilliseconds: 10000,
},
],
})) {
products.push(product);
}

All full-text conditions combined into one query must use the same freshness values. Harper returns 400 when combined conditions specify different values.

A non-waiting HTTP query can return Harper-Index-Coverage with the admitted state, lag upper bound, and requested tolerance. A zero-size page performs no native search and carries no coverage proof. A waiting Table.search() establishes coverage as its iterator is consumed. A search_by_conditions request completes only after its wait and search finish. Waiting queries do not emit a coverage header before completion.

If acceptable coverage is not reached before waitForIndexMilliseconds expires, Harper returns 503 with code: "DERIVED_INDEX_LAGGING" and retryable: true. For Table.search(), the error is raised while the async iterator is consumed. An HTTP response may already be streaming, so clients should also handle a terminal stream error rather than relying only on the initial status.

The same code is also used when prolonged derived-index lag causes Harper to reject local writes. Query waits can be retried according to the caller's freshness needs; rejected writes should use backoff while the index catches up. See Write backpressure.

Prefix result window​

matches_prefix and matches_fuzzy_prefix are autocomplete-style record searches. Any expression containing one of these modes uses a 100-record native result window, and offset + limit cannot exceed that window. If an unbounded query has more than 100 matches, Harper returns 400 and requires a limit instead of silently truncating the result. These modes return matching records, not a separate list of suggested terms.

Other match modes page through the native result window as needed.

REST​

Use the comparator in the collection query string:

GET /Product/?catalogSearch=matches=waterproof%20trail&limit(20)
GET /Product/?catalogSearch=matches_phrase=trail%20running&select(id,name,$score,$highlights)
GET /Product/?catalogSearch=matches_prefix=waterproof%20tra&limit(10)

REST supports all positive and negated full-text comparators. It can express the query text, selected fields, and pagination.

Repeat the index parameter to combine the required positive and negated conditions:

GET /Product/?catalogSearch=matches=waterproof&catalogSearch=not_matches=leather&limit(20)

REST can return configured highlights by selecting $highlights. Use Table.search() or search_by_conditions for explicit fields, condition-level includeHighlights, or freshness controls.

Operations API​

search_by_conditions uses the same condition shape:

{
"operation": "search_by_conditions",
"database": "catalog",
"table": "Product",
"limit": 20,
"get_attributes": ["id", "name", "$score", "$highlights"],
"conditions": [
{
"attribute": "catalogSearch",
"comparator": "matches_all",
"value": "waterproof trail",
"fields": ["name", "description"],
"includeHighlights": true,
"maxIndexLagMilliseconds": 0,
"waitForIndexMilliseconds": 10000
}
]
}

If an index is unknown, needs-rebuild, or rebuilding, the query returns 503 with code: "INDEX_REBUILDING" and retryable: true. An index in terminal unavailable state returns a generic, non-retryable 503; inspect readiness and logs instead of retrying it as a rebuild. Other busy paths can also return a generic 503, so clients should branch on the code rather than the status alone. Unreadable source fields return 403. Invalid declarations, unsupported match modes, incompatible combinations, and out-of-range options return 400 responses.