Solr Glossary

These are common terms used with Solr.

Solr Terms

Where possible, terms are linked to relevant parts of the Solr Reference Guide for more information.

Jump to a letter:

A B C D E F G H I J K L M N O P Q R S T U V W X Y Z

A

Atomic updates

An approach to updating only one or more fields of a document, instead of reindexing the entire document.

B

Boolean operators

These control the inclusion or exclusion of keywords in a query by using operators such as AND, OR, and NOT.

C

Cluster

In Solr, a cluster is a set of Solr nodes operating in coordination with each other via ZooKeeper, and managed as a unit. A cluster may contain many collections. See also Solr Cluster Types and SolrCloud.

Collection

The complete logical set of searchable documents that share a schema and configuration.

In SolrCloud, a collection may be divided up into multiple logical shards, which may in turn be distributed across many nodes for scalability and fault tolerance. Each collection encompasses all the shards and their replicas.

Standalone installations and user-managed clusters do not manage collections as first-class entities; instead they work directly with individual cores.

Commit

To make document changes permanent in the index. In the case of added documents, they would be searchable after a commit.

Core

In Solr’s implementation, a core is the physical instance that represents a Replica. When you create a replica, Solr creates a core to represent it. Multiple cores can be hosted on a single Node. Each core maintains exactly one Lucene Index on disk.

Historically the term "core" has mostly been used as a synonym for replica, but the term "core" can be confusing because in everyday English it implies something central and singular. Since there may be many replicas in Solr, and they are distributed across the cluster "Replica" is the preferred term. Core is mostly only used for historical reasons in the code base and other places where renaming things would be disruptive.

See also SolrCloud.

Corpus

The set of documents available for indexing, irrespective of whether or not they are currently indexed in Solr.

Core reload

To re-initialize a Solr core after changes to the schema file, solrconfig.xml or other configuration files.

D

Distributed search

Distributed search is one where queries are processed across more than one Shard.

Document

A group of fields and their values. Documents are the basic unit of data in a collection. Documents are assigned to shards using standard hashing, or by specifically assigning a shard within the document ID. Documents are versioned after each write operation.

E

Ensemble

A ZooKeeper term to indicate multiple ZooKeeper instances running simultaneously and in coordination with each other for fault tolerance.

F

Facet

The arrangement of search results into categories based on indexed terms.

Field

The content to be indexed/searched along with metadata defining how the content should be processed by Solr.

Follower

A Replica that is not the Leader for its Shard. Follower replicas receive index updates from the leader replica and serve queries. See also SolrCloud.

I

Inverse document frequency (IDF)

A measure of the general importance of a term. It is calculated as the number of total Documents divided by the number of Documents that a particular word occurs in the collection. See http://en.wikipedia.org/wiki/Tf-idf and the Lucene TFIDFSimilarity javadocs for more info on TF-IDF based scoring and Lucene scoring in particular. See also Term frequency.

Index

The physical data structures written to disk by Apache Lucene. Each Core (Replica) maintains exactly one Lucene index on disk, containing the actual inverted indexes, stored fields, and other data structures that enable search.

Inverted index

A way of creating a searchable index that lists every word and the documents that contain those words, similar to an index in the back of a book which lists words and the pages on which they can be found. When performing keyword searches, this method is considered more efficient than the alternative, which would be to create a list of documents paired with every word used in each document. Since users search using terms they expect to be in documents, finding the term before the document saves processing resources and time.

L

Leader

A single Replica for each Shard that serves as the source-of-truth and coordinates index updates (document additions or deletions) to the follower replicas in the same shard. This is a transient responsibility assigned to a replica via an election; if the current leader goes down, another replica will automatically be elected to take its place. See also SolrCloud.

M

Metadata

Literally, data about data. Metadata is information about a document, such as its title, author, or location.

N

Natural language query

A search that is entered as a user would normally speak or write, as in, "What is aspirin?"

Node

An instance of a running Solr process that services search and indexing requests. A node is a JVM instance running Solr on a Server.

O

Optimistic concurrency

Also known as "optimistic locking", this is an approach that allows for updates to documents currently in the index while retaining locking or version control.

Overseer

A single node in SolrCloud that is responsible for processing and coordinating actions involving the entire cluster. It keeps track of the state of existing nodes, collections, shards, and replicas, and assigns new replicas to nodes. This is a transient responsibility assigned to a node via an election, if the current Overseer goes down, a new node will be automatically elected to take its place. See also SolrCloud.

Q

Query parser

A query parser processes the terms entered by a user.

R

Recall

The ability of a search engine to retrieve all of the possible matches to a user’s query.

Relevance

The appropriateness of a document to the search conducted by the user.

Replica

The physical manifestation of a logical Shard. A replica is the actual running instance that holds and serves the documents belonging to that shard. A shard must have at least one replica to exist physically, and may have multiple replicas for redundancy and fault tolerance. All replicas of the same shard contain the same subset of documents. See also SolrCloud.

Replication

A method of copying a leader index from one server to one or more "follower" or "child" servers.

RequestHandler

Logic and configuration parameters that tell Solr how to handle incoming "requests", whether the requests are to return search results, to index documents, or to handle other custom situations.

S

SearchComponent

Logic and configuration parameters used by request handlers to process query requests. Examples of search components include faceting, highlighting, and "more like this" functionality.

Server

The hardware or virtual machine that hosts Solr software. A server may run one or more Solr Nodes.

Shard

A logical slice of a Collection. Each shard represents a partition containing a subset of the collection’s documents. A shard exists physically as one or more Replicas, which may be distributed across multiple Nodes for fault tolerance and scalability. See also SolrCloud.

SolrCloud

Umbrella term for a suite of functionality in Solr which allows managing a Cluster of Solr Nodes for scalability, fault tolerance, and high availability.

Solr Schema (managed-schema.xml or schema.xml)

The Solr index Schema defines the fields to be indexed and the type for the field (text, integers, etc.). By default schema data can be "managed" at run time using the Schema API and is typically kept in a file named managed-schema.xml which Solr modifies as needed, but a collection may be configured to use a static Schema, which is only loaded on startup from a human edited configuration file - typically named schema.xml. See Schema Factory Configuration for details.

SolrConfig (solrconfig.xml)

The Apache Solr configuration file. Defines indexing options, RequestHandlers, highlighting, spellchecking and various other configurations. The file, solrconfig.xml, is located in the Solr home conf directory.

Spell Check

The ability to suggest alternative spellings of search terms to a user, as a check against spelling errors causing few or zero results.

Stopwords

Generally, words that have little meaning to a user’s search but which may have been entered as part of a natural language query. Stopwords are generally very small pronouns, conjunctions and prepositions (such as, "the", "with", or "and")

Suggester

Functionality in Solr that provides the ability to suggest possible query terms to users as they type.

Synonyms

Synonyms generally are terms which are near to each other in meaning and may substitute for one another. In a search engine implementation, synonyms may be abbreviations as well as words, or terms that are not consistently hyphenated. Examples of synonyms in this context would be "Inc." and "Incorporated" or "iPod" and "i-pod".

Standalone

An informal term referring to Solr nodes that do not utilize Apache Zookeeper and thus do not provide the centralized configuration management that is available in SolrCloud mode. This includes both standalone installations and User-Managed clusters. In source code and documentation, "Standalone" may refer to either User-Managed Mode or standalone deployments. See also Solr Cluster Types and Cluster.

T

Term frequency

The number of times a word occurs in a given document. See http://en.wikipedia.org/wiki/Tf-idf and the Lucene TFIDFSimilarity javadocs for more info on TF-IDF based scoring and Lucene scoring in particular. See also Inverse document frequency (IDF).

Transaction log

An append-only log of write operations maintained by each Replica. This log is required with SolrCloud implementations and is created and managed automatically by Solr.

U

User-Managed Cluster

A mode of operating a Solr Cluster without the centralized coordination provided by ZooKeeper in SolrCloud mode. In user-managed mode, cluster coordination activities must be performed manually or with local scripts. This includes shard creation, document routing, leader/follower configuration, and load balancing. Also known as Standalone mode. See also Solr Cluster Types.

W

Wildcard

A wildcard allows a substitution of one or more letters of a word to account for possible variations in spelling or tenses.

Z

ZooKeeper

Also known as Apache ZooKeeper. The system used by SolrCloud to keep track of configuration files and node names for a cluster. A ZooKeeper cluster is used as the central configuration store for the cluster, a coordinator for operations requiring distributed synchronization, and the system of record for cluster topology. See also SolrCloud.