Healthchecks

Solr provides two distinct tools for checking whether things are working, and they answer different questions. Confusing them is easy, since both are casually called a "healthcheck", so this page describes each and helps you pick the right one.

Node Health Endpoint bin/solr healthcheck command

Answers

"Is this node alive and part of the cluster?"

"Does this collection actually have data, and is every replica in sync?"

Scope

A single node

A single collection, across all its shards and replicas

Checks the index/data?

No — metadata only

Yes — runs real queries

Interface

HTTP endpoint

Command-line tool

Typical use

Load balancer / orchestrator liveness or readiness probe

Manual or scripted diagnosis after deployment, restart, or an incident

Node Healthcheck

The api/node/health endpoint (NodeHealth) reports whether a single Solr node is alive and able to participate in the cluster. It is designed to be cheap enough to poll frequently from a load balancer or an orchestrator like Kubernetes.

None of these checks touch the index. A node with a collection that has zero documents, or a collection whose query results are wrong, will still report a healthy 200 OK here. If you need to confirm a collection actually has data, see Collection Healthcheck below.

What the endpoint checks — and how it fails — depends on whether the node is running in SolrCloud mode or user-managed (leader-follower) mode.

SolrCloud Mode

In SolrCloud mode, the endpoint returns HTTP 200 OK only if all of the following are true, and HTTP 503 Unavailable otherwise:

  • The node’s core container has finished starting up.

  • The node is connected to ZooKeeper.

  • The node is listed in ZooKeeper’s live_nodes.

curl "http://localhost:8983/api/node/health"

Healthy node:

{
  "responseHeader":{"status":0,"QTime":1},
  "status":"OK"
}

Unhealthy node (HTTP 503):

{
  "responseHeader":{"status":503,"QTime":1},
  "status":"FAILURE",
  "error":{"msg":"Not connected to ZK"}
}

Rolling Restarts: requireHealthyCores

Add requireHealthyCores=true to additionally require that every local replica belonging to an active shard has finished initializing — i.e., none are in the RECOVERING or DOWN state. This is useful as a readiness probe during a rolling restart, so an orchestrator doesn’t move on to the next node while the one it just restarted still has replicas recovering.

curl "http://localhost:8983/api/node/health?requireHealthyCores=true"

User-Managed (Leader-Follower) Mode

In user-managed replication, the endpoint instead checks how far behind each local follower core is from its leader, in Lucene commit generations. Set maxGenerationLag=<n> to fail the healthcheck once a follower falls more than <n> generations behind; without it, the check simply reports OK once a follower has replicated at least once, even if it later falls arbitrarily far behind.

curl "http://localhost:8983/api/node/health?maxGenerationLag=100"

Healthy node (all followers within the allowed lag):

{
  "responseHeader":{"status":0,"QTime":1},
  "status":"OK"
}

Unhealthy node (a follower has fallen too far behind, HTTP 503):

{
  "responseHeader":{"status":503,"QTime":1},
  "status":"FAILURE",
  "error":{"msg":"Cores violating maxGenerationLag:100.\nCore collection1 is lagging by 137 generations"}
}

Collection Healthcheck

While the Node Health endpoint checks process liveness, the bin/solr healthcheck tool verifies whether a specific collection is functioning properly. It queries every replica directly (non-distributed) to compare document counts, confirm each shard has a leader, and verify all replicas are in the ACTIVE state.

Because it executes real queries against individual replicas, it is a heavier operation than the Node Health endpoint and is meant as an on-demand CLI tool rather than a continuously polled HTTP healthcheck for load balancers. Use it after deployments, restarts, or incidents to confirm collection integrity.

bin/solr healthcheck -c gettingstarted
{
  "collection":"gettingstarted",
  "status":"healthy",
  "numDocs":42,
  "numShards":2,
  "shards":[
    {
      "shard":"shard1",
      "status":"healthy",
      "replicas":[
        {"name":"core_node1", "url":"...", "numDocs":21, "status":"active", "leader":true}
      ]
    }
  ]
}
Healthy, well-functioning replicas often disagree on exact document counts due to slightly staggered autoCommit periods or in-flight updates. Small, momentary count discrepancies between replicas are normal and should not automatically be treated as an issue unless the mismatch persists.

See the healthcheck command reference for the full list of options.