Opensearch Migration Reference

Explanatory material for the OpenSearch Migration Guide, which is the procedure this reference supports.

Nothing here is required in order to follow that procedure — the guide links to each section at the point where it matters. Read a section when you want to know why a step is written the way it is, or when a step raises a question the procedure does not answer.

The Four Positions#


The position is set with DOT_FEATURE_FLAG_OPEN_SEARCH_PHASE and read at startup. An absent or unrecognised value means position 0.

Position 0: Migration Not Started#

Elasticsearch only, for both reads and writes. The starting state, and the state dotCMS returns to if startup validation fails at position 1 or 2.

Position 1: Dual-Write, ES Reads#

Writes go to both engines. Reads come from Elasticsearch.

OpenSearch is the shadow engine here. A write that fails to reach it is logged as a warning and otherwise ignored — deliberately, so that a problem with the new engine cannot break a working site. The cost of that design is covered under Startup Validation: the write is not retried and not queued, so a failed shadow write is a permanent gap until the next full reindex.

Nothing visible to a visitor changes at position 1. If OpenSearch goes down entirely, your site continues to work.

Position 2: Dual-Write, OS Reads#

Writes still go to both engines. Reads now come from OpenSearch.

The relationship inverts: OpenSearch is now primary and Elasticsearch is the shadow. That has two consequences worth knowing before you get here.

Search results now come from the new engine, which is the point — this is where you find out whether it serves your site correctly. And an OpenSearch outage is now visible: searches fail, and publishing content fails with an error naming the unreachable host.

Elasticsearch remains complete and current throughout, so returning to position 1 costs nothing.

Position 3: OpenSearch Only#

OpenSearch only. Elasticsearch is neither written to nor read from, and begins going stale immediately.

The failure behaviour changes here, and this is the single most important thing to carry away from this section. At positions 1 and 2, if dotCMS cannot validate OpenSearch at startup, it falls back to Elasticsearch and keeps running. At position 3 it will not do that — Elasticsearch is no longer being maintained, so falling back to it would mean serving stale content and writing new content into an engine that is being decommissioned. Instead dotCMS retries the connection and then stops.

This is correct behaviour and it is also a new operational obligation. See Startup Validation.

Why OpenSearch 3.x#

dotCMS checks the version at startup and refuses to migrate to anything else. The neutral search client it uses to talk to OpenSearch 3.x cannot address a 1.x cluster the same way, so a 1.x target would not be a slower migration — it would be a broken one. The check exists to catch a mis-pointed endpoint before it does damage.


The Readiness Report#


curl -su "$DOT_AUTH" "$DOT/api/v1/index/migration/readiness" | jq

phase#

{
  "current": 2,
  "dualWrite": true,
  "name": "PHASE_2_DUAL_WRITE_OS_READS",
  "readEngine": "OpenSearch",
  "writeEngines": ["Elasticsearch", "OpenSearch"]
}

current is the runtime position, not the configured one. That distinction is the whole reason this endpoint is the way to confirm a position change: the environment variable tells you what you asked for, and this tells you what you got.

readEngine is the fastest way to know whether a test you are about to run is meaningful.

verdict#

{
  "blockers": [],
  "outOfSyncCount": 0,
  "safeToAdvance": true,
  "safeToRollback": true,
  "summary": "..."
}

safeToAdvance is phase-aware — it answers "is it safe to move to the next position from where I am now," not a general health question.

safeToRollback means moving back will not lose data. It does not mean rolling back will fix whatever the blockers are complaining about. Those are separate questions and the report answers only the first. If you are looking at blockers recommending a reindex, the reindex is the remedy; rolling back is not a lighter alternative to it.

outOfSyncCount counts indexes, not documents. Two indexes out of sync by one document each and two indexes out of sync by a million documents each both report 2. Read the blockers for what is actually wrong.

Per-Index Detail#

Each index appears with its name, its state on each engine, and a verdict:

{
  "indexName": "sitesearch_20260925041401_896e0f61-...",
  "es": { "exists": true, "docCount": 8, "alias": "new", "physicalName": "cluster_….sitesearch_…" },
  "os": { "exists": true, "docCount": 8, "alias": "new", "physicalName": "cluster_….sitesearch_….os" },
  "driftPercent": 0.0,
  "verdict": "IN_SYNC",
  "recommendation": "In sync — no action needed."
}

physicalName is the actual index on each engine. OpenSearch copies carry a .os suffix; this is expected and is not drift.

Content indexes appear under their own keys (CONTENT_LIVE, CONTENT_WORKING), and Site Search indexes under siteSearch.

What It Cannot See#

It matches indexes by name. If OpenSearch holds an equivalent index under a different name, the report will say the expected one is missing, because from where it is standing that is true. The advice it gives in that case — run a full reindex — is correct, because a reindex makes the names correspond. But the blocker is telling you about names, not necessarily about content.

It compares document counts, not documents. Matching counts are good evidence of parity and are not proof of it.

It cannot answer once Elasticsearch is gone. After you decommission the old engine the report has only one side to look at. Do not plan to use it as a long-term monitoring tool.


TLS and Certificates#


If curl against your OpenSearch cluster works with -k and fails without it, dotCMS will hit the same wall. Two ways to deal with it.

Install the certificate properly. Add your CA or self-signed certificate to the trust store the dotCMS JVM uses. This is the right answer for anything that will outlive the migration, and it means no dotCMS-specific TLS settings at all.

Or tell dotCMS to accept it:

OS_TLS_ENABLED=true
OS_TLS_TRUST_SELF_SIGNED=true

Both default to false. The full set is OS_TLS_ENABLED, OS_TLS_TRUST_SELF_SIGNED, OS_TLS_CERT_REQUIRED, OS_TLS_CA_CERT, OS_TLS_CLIENT_CERT and OS_TLS_CLIENT_KEY. As environment variables they take dotCMS's usual DOT_ prefix.

OS_TLS_ENABLED falls back to ES_TLS_ENABLED if it is not set, so an installation that already configured TLS for Elasticsearch inherits it. The trust settings have no such fallback — they must be set explicitly.

Two things that catch people:

  • An https:// endpoint with TLS disabled will not work. The endpoint scheme and the TLS setting have to agree. If DOT_OS_ENDPOINTS starts with https://, TLS must be on.
  • Trusting a self-signed certificate is a separate setting from enabling TLS. Enabling TLS against a self-signed certificate without also trusting it fails, and the error looks like a connection problem rather than a certificate one. The trust settings are read regardless of OS_TLS_ENABLED, so they apply even when TLS comes on from the endpoint scheme — which makes omitting them the more damaging of the two mistakes.

The Service Account#


dotCMS needs an OpenSearch account that can manage and search its own indexes, plus one cluster-level permission.

cluster:monitor/main#

This is the one that matters and the one that is easy to miss.

dotCMS uses it to ask the cluster to identify itself — the request behind the version check at startup. Without it, that request returns 403, startup validation fails, and at positions 1 and 2 dotCMS returns to position 0 and carries on serving from Elasticsearch.

Nothing about your site looks wrong when this happens. The migration simply is not running. Prove this permission works at 2.4, under the service account's own credentials, before you rely on it.

Full Permission Set#

dotCMS does not need an OpenSearch administrator account. Create a user — conventionally dotcms-es-user — mapped to a least-privilege role scoped to your own indexes.

Cluster permissions:

cluster:monitor/health
cluster:monitor/state
cluster:monitor/nodes/stats
cluster:monitor/main
indices:data/write/bulk
indices:data/read/scroll
indices:data/read/scroll/clear

Index permissions on cluster_<your-cluster-name>*: indices_all, indices_monitor

Index permissions on * (read-only, for listing and stats): indices:monitor/stats, indices:monitor/settings/get, indices:admin/aliases/get

indices_all on your own index pattern already expands to everything dotCMS needs — it creates, rotates and deletes indexes as part of normal operation, so read and write alone are not sufficient.

If you already have an Elasticsearch service account, the only thing it is missing is cluster:monitor/main. Everything else in the historic dotCMS role carries over unchanged.

What Not to Do#

Do not use the cluster's administrative account. It works, which is the problem — it also masks permission mistakes that will surface later under a correctly-scoped account, and it leaves an over-privileged credential in your configuration after the migration is over.


Templates and Viewtools#


Why They Are Dangerous#

DeprecatedReplacement
$estool.esSearch($query)$estool.search($query)
$estool.esRaw($query)$estool.raw($query)

These two do not follow the position setting. They read Elasticsearch directly regardless of which engine the rest of your installation is using.

Through positions 0, 1 and 2 that is invisible. Both engines hold the same content, so a template calling esSearch returns exactly what it always did — correct results, no error, nothing in the log. Position 2 is the one window where it can actually diverge: your site reads OpenSearch while these methods read Elasticsearch, so a write that reached one engine and not the other shows up as a page disagreeing with the rest of the site. Rare, and silent when it happens.

At position 3, what happens depends on your release.

From 26.09.23-01 these methods refuse to run. Elasticsearch no longer receives writes there, so rather than answer from an index that is going stale, the API throws DotStateException — which fails the page render for a template, and surfaces as an exception for a plugin. Support can enable FEATURE_FLAG_OPEN_SEARCH_LEGACY_ES_SEARCH_RETURNS_NULL to make them return nothing and log a rate-limited warning instead, so a page renders without those results rather than failing outright. That flag is a stopgap for an installation caught mid-conversion, not a way to keep the methods.

On earlier releases they keep answering from the frozen Elasticsearch copy. Results go quietly out of date with no error, and the failure arrives later, when you decommission the old engine — weeks after the migration, looking like an unrelated incident.

Either way you cannot find these by testing before position 3. You find them by grepping, which is Step 1. On a current release, unfixed code takes pages down at cutover; on an older one it serves stale results until the hardware is switched off. The first is louder and considerably kinder.

What Changed in 26.05.19-01#

If you are on 26.05.19-01 or later, this has already happened to your installation.

Most templates are unaffected. #foreach, $results.size(), $item.title, $results.aggregations, $results.hits and $results.response all continue to resolve, and the structure of the aggregation tree is unchanged. Three things did change.

$item.map.fieldName no longer resolves. Results from search() are ContentMap objects, which have no map property. Read fields directly — $item.fieldName or $item.get("fieldName"). The old expression returns nothing and the value disappears from the page without an error.

$raw.took is now $raw.tookInMillis, a plain number of milliseconds.

$raw.toString() no longer produces JSON. It returns a Java object description. Anything downstream expecting JSON receives something else, without error.

Producing JSON#

Use the $json viewtool:

#set($raw = $estool.raw($q))
<script>
  var data = $json.generate($raw);
</script>

This produces dotCMS's neutral response shape, which uses different keys from the ones Elasticsearch emitted. If a script was parsing the old output, it needs remapping:

Elasticsearch emittedThe neutral response emits
tooktookInMillis
hits.totalhits.totalHits
_indexindex
_idid
_scorescore
_sourcesourceAsMap

timed_out, _shards and max_score have no equivalent, and totalHits carries value without the relation qualifier.

Elasticsearch's own wire format is not available under any method. If a consumer requires it specifically, that consumer needs restructuring rather than a different call.

If you only need part of the response, build the object you want and generate that:

#set($out = {"total": $raw.hits.totalHits.value, "items": []})
#foreach($hit in $raw.hits.hits)
  #set($ignore = $out.items.add($hit.sourceAsMap))
#end
<script>var data = $json.generate($out);</script>

The Field-Shadowing Trap#

esSearch() returns raw content objects rather than ContentMap, which is why $item.map.fieldName works there. It also exposes a problem that search() does not: a custom field whose name matches a built-in getter — contentTypeId, for example — resolves to the getter rather than to your field's value. Moving to search() fixes it. If a converted template suddenly shows a different value where it previously showed a wrong one, this is why.

Testing a Conversion#

Replace quiet references ($!{...}) with ${...} while you test. The quiet form prints nothing when an expression fails to resolve, so a template that has lost a value renders as a page with a blank spot rather than as an error — which is exactly the failure mode all three of the changes above produce.


Plugins#


Three levels of involvement, in ascending order of work.

Level 1: Method Renames#

// before
APILocator.getContentletAPI().esSearch(query, live, user, respectFrontendRoles);
// after
APILocator.getContentletAPI().search(query, live, user, respectFrontendRoles);
Stop usingUse instead
ContentletAPI.esSearch(...)ContentletAPI.search(...)
ContentletAPI.esSearchRaw(...)ContentletAPI.searchRaw(...)
APILocator.getEsSearchAPI()APILocator.getSearchAPI()
ContentletAPIPreHook.esSearch / esSearchRawContentletAPIPreHook.search / searchRaw
ContentletAPIPostHook.esSearch / esSearchRawContentletAPIPostHook.search / searchRaw

The replacements are default interface methods, so existing implementors inherit them without change. The hook replacements have no-op default bodies — override them only if you actually intercept search.

Level 2: Return Types#

WasIs now
ESSearchResultsContentSearchResults<Contentlet>
org.elasticsearch.action.search.SearchResponseContentSearchResponse
ElasticsearchExceptionDotSearchException

ContentSearchResults<T> is a typed List<T>, so the (Contentlet) casts go away. DotSearchException extends DotRuntimeException and so remains unchecked — only code that explicitly catches the Elasticsearch type needs an edit.

Watch the imports. IndexStats, ClusterStats and NodeStats each exist in more than one package. Use the ones in com.dotcms.content.index.domain. An older import compiles cleanly and fails at runtime, which is the worst of both.

Level 3: Index Metadata#

Separately deprecated, and the only group where the replacements are not mechanical.

Stop usingUse instead
IndiciesAPIVersionedIndicesAPI
APILocator.getIndiciesAPI()APILocator.getVersionedIndicesAPI()
IndiciesInfoVersionedIndices
IndiciesFactoryIndicesFactory
ContentletIndexAPI.timestampFormatterContentletIndexAPI.threadSafeTimestampFormatter

Note the spelling correction — Indicies becomes Indices.

Three are not drop-in replacements:

  • loadIndicies() becomes loadIndices(String version) — it takes a version argument, returns an Optional, and ignores legacy rows whose version is null. A plugin that assumed a single global index set has to decide which version it wants.
  • loadIndicies(Connection) splits into loadAllIndices(), loadNonVersionedIndices() and loadDefaultVersionedIndices(). There is no connection-scoped variant.
  • point(IndiciesInfo) becomes saveIndices(VersionedIndices), which throws DotDataException when the object carries no version. The old method accepted anything, so this is the most likely upgrade break in the group.

The timestampFormatter replacement is worth taking regardless of the migration: the old constant is a SimpleDateFormat, which is not thread-safe.

A Note on Scope#

Everything in com.dotcms.content.index.*, the engine-specific implementations, and anything reached through CacheLocator or FactoryLocator exists so that dotCMS can swap search engines. It is not an extension surface and it will keep moving. Plugins should stay on APILocator and the published interfaces.


Queries Are JSON#


Every method on $estool takes a raw engine JSON query body — the same syntax you would send to the search engine itself. It does not accept Lucene query strings.

## Correct
#set($q = '{"query":{"bool":{"filter":[{"term":{"live":true}}]}},"size":20}')

## Not accepted — throws "Unable to parse the given query"
#set($q = '+contentType:News +live:true')

If you want Lucene syntax, that is $dotcontent.pull(), which is a different tool and unaffected by this migration.

This catches people because the error arrives at runtime rather than at save time, and because plenty of dotCMS documentation elsewhere uses Lucene syntax for other purposes. You can test a query in Dev Tools → ES Search before putting it in a template.


Query Normalization#


$estool.search() and $estool.raw() lowercase the entire query before executing it. This is deliberate: it means a mixed-case field name like contentType still resolves to the physical index field contenttype.

Two consequences.

Values are lowercased too, not just field names. Neither method supports a case-sensitive exact match on a value.

Aggregation names come back lowercased. An aggregation you declared as "tagAgg" is keyed tagagg in the response:

#set($q = '{"aggs":{"tagagg":{"terms":{"field":"tags","size":10}}},"size":0}')
#set($results = $estool.search($q))

## remember: aggregation names come back lowercased
#set($tags = $results.aggregations.get("tagagg"))

#foreach($bucket in $tags.buckets)
  $bucket.keyAsString ($bucket.docCount)
#end

Looking the aggregation up under the name you wrote returns null, and a #foreach over null renders nothing rather than raising an error — so this presents as an empty section of a page rather than as a failure. It is the second most common surprise when converting a template, after $item.map.

Nested aggregations and top_hits are preserved, so templates that walk deeper into the tree continue to work unchanged.

The deprecated esRaw() normalizes identically — its own documentation says the query is lowercased exactly as raw() does it. So this is a standing property of the viewtool, not something a template meets for the first time when it switches methods. If your aggregation names already work, they will keep working.


Indexes and Reindexing#


Index Naming#

cluster_<cluster-name>.live_<timestamp>
cluster_<cluster-name>.working_<timestamp>
cluster_<cluster-name>.sitesearch_<timestamp>_<uuid>

On OpenSearch the same index carries a .os suffix. So the OpenSearch copy of cluster_prod.live_20260924235216 is cluster_prod.live_20260924235216.os. Seeing the suffix in the readiness report or in _cat/indices is expected and is not a sign of drift.

The timestamp is a generation. dotCMS does not edit an index in place when it rebuilds — it creates a new one with a fresh timestamp and moves the alias. Your cluster therefore normally holds more than one generation, and the alias is what says which is live.

What a Reindex Does#

curl -su "$DOT_AUTH" -X POST "$DOT/api/v1/esindex/reindex"
curl -su "$DOT_AUTH" "$DOT/api/v1/esindex/reindex" | jq       # progress
curl -su "$DOT_AUTH" -X DELETE "$DOT/api/v1/esindex/reindex"  # cancel

It walks all content and writes it into fresh indexes on every engine currently being written to, then repoints the aliases. It runs in the background; the site stays up.

At position 1 or 2 that means both engines get a new generation and end up corresponding. This is why a reindex is the standard remedy for almost everything the readiness report complains about — it does not repair a divergence so much as replace both sides with a matched pair.

When You Need One#

  • Once at position 1, always — Rule 1. Dual-write only carries changes made after you switch it on; the reindex is what moves everything that already existed.
  • After a shadow-write warning. A failed shadow write is not retried and not queued. A reindex is the only repair.
  • Whenever the readiness report says so. Its recommendation field says it directly.

What It Does Not Do#

It does not fix code. If a template or plugin is reading the wrong engine because it calls a deprecated method, no amount of reindexing changes that — see Templates and Viewtools and Plugins.


Site Search#


Site Search indexes are separate from content indexes, and they behave differently enough to be worth their own section.

Tracked and Migrated#

Site Search indexes appear in the readiness report under siteSearch, with the same per-engine detail as content indexes — document counts on both sides, drift percentage, verdict and recommendation. Treat them exactly as you treat the content indexes when deciding whether it is safe to advance.

They survive the cutover to position 3 in sync, and dual-write creates them correctly on both engines when a crawl runs mid-migration.

Built by Crawling#

Publishing content does not update a Site Search index. Only a crawl does. So a Site Search index will sit at the same document count through any amount of content editing, and that is correct rather than a sign of drift.

The practical consequence: if you publish content during the migration and want Site Search to know about it, you have to run a crawl. Which brings us to the part that needs care.

When a Crawl Rebuilds#

This is the thing to know before you touch the Site Search tool during a migration. Ticking Incremental is a request, not a guarantee: dotCMS runs the crawl in place only when every one of these holds, and otherwise rebuilds the index from scratch.

ConditionWhy a rebuild happens without it
The job is scheduled, not Run: NowThe UI disables Incremental for on-demand runs, so these are always full crawls
The job has run beforeAn incremental needs a previous run's timestamp to define "changed since"
The index already exists and is not emptyThere is nothing to add to
The index exists on every write engine, with matching document countsSee below — this one is specific to migrating

So a newly created job rebuilds on its first run even with Incremental ticked, and runs in place from the second onward.

The mirror condition is the one that concerns you during a migration. While dual-write is on, a Site Search index is one logical index copied across both engines. If the copy is missing on one engine, or the two copies have drifted apart, an in-place write would make things worse — it would either auto-create a twin with the wrong mapping, or layer new documents on top of two already divergent copies and never reconcile them. So dotCMS deliberately forces a full rebuild instead, which recreates matching copies on every engine and repoints the alias, healing the mirror on that crawl. It says so in the log:

INFO Site-search index `<name>` is missing on a write engine or its engine copies are out of sync;
     forcing a full rebuild instead of an incremental crawl to restore the mirror.

If you see that line, the crawl did its job — but something had already gone wrong with dual-write for that index, and it is worth knowing what.

Rebuilds Are Not Reversible#

When the alias moves to the new index, the previous index is removed from both engines. Nothing is orphaned — which is good for housekeeping — but it also means there is no previous Site Search index to fall back to if the rebuild produces a worse result than what it replaced.

If your Site Search configuration is complicated, or if the crawl scope has changed since the index was last built, satisfy yourself about the configuration before triggering a crawl mid-migration rather than after.

Practical Advice#

Run your Site Search crawl before Step 3, so the index is built and stable before dual-write starts. Then check it in the readiness report at 3.5, 4.1 and 5.1 alongside everything else, and avoid re-crawling during the migration unless you have a reason.

Finally, note that an incremental crawl only covers the window since the job's previous run. It will not pick up anything that changed before that, so a job whose scope or configuration has changed needs a deliberate full crawl rather than an incremental one.


Troubleshooting#


SymptomLikely causeWhere to go
Position change appears to do nothingStartup validation failed; dotCMS returned to position 0Startup Validation, and 3.3 / 4.3
Site looks perfect at position 2Possibly reading Elasticsearch because the position did not take — check readEngine4.3
A page's results went empty after conversion$item.map.field, or a mixed-case aggregation nameTemplates and Viewtools, Query Normalization
JavaScript on a page broke after conversion$raw.toString() was feeding it JSONTemplates and Viewtools
"Unable to parse the given query"Lucene syntax where JSON is requiredQueries Are JSON
WARN … OS shadow write failedA document reached one engine and not the other, permanently3.5, Indexes and Reindexing
Publishing fails with your OpenSearch hostname in the messageOpenSearch unreachable at position 2Startup Validation
dotCMS will not start after moving to position 3OpenSearch unreachable or unauthorised; this is correct behaviourStartup Validation
403 on the readiness endpointYour account is not in the readiness role3.1
Readiness says an OpenSearch index is "missing"Names do not correspond; a reindex makes themThe Readiness Report, Indexes and Reindexing
A page or plugin fails at position 3 naming the deprecated pathDeprecated methods in your own code; from 26.09.23-01 they refuse to run thereTemplates and Viewtools, Plugins
Search worked before the migration, fails after decommissioningSame cause, on a release earlier than 26.09.23-01Templates and Viewtools, Plugins
Results ordered differently but completeUsually normalization; usually harmlessQuery Normalization

Command Reference#


Ask an engine its version

curl -sk -u "$OS_AUTH" "$OS_URL/" | jq .version.number      # new engine
curl -sk -u "$SRC_AUTH" "$SRC_URL/" | jq .version.number    # current engine

List an engine's indexes

curl -sk -u "$OS_AUTH" "$OS_URL/_cat/indices?v"
curl -sk -u "$SRC_AUTH" "$SRC_URL/_cat/indices?v"

Read the readiness report

curl -su "$DOT_AUTH" "$DOT/api/v1/index/migration/readiness" | jq
curl -su "$DOT_AUTH" "$DOT/api/v1/index/migration/readiness" | jq '.phase'
curl -su "$DOT_AUTH" "$DOT/api/v1/index/migration/readiness" | jq '.verdict'
curl -su "$DOT_AUTH" "$DOT/api/v1/index/migration/readiness" | jq '.siteSearch'

Control a reindex

curl -su "$DOT_AUTH" -X POST   "$DOT/api/v1/esindex/reindex"   # start
curl -su "$DOT_AUTH"           "$DOT/api/v1/esindex/reindex"   # progress
curl -su "$DOT_AUTH" -X DELETE "$DOT/api/v1/esindex/reindex"   # cancel

Change the position

DOT_FEATURE_FLAG_OPEN_SEARCH_PHASE=0|1|2|3

On every node, then restart, then confirm with the .phase command above.

Find affected code

grep -rn "esSearch\|esRaw" <path-to-your-vtl>
grep -rn "esSearch\|esRaw" <dotcms>/application/apivtl/
grep -rln "esSearch\|esRaw" <dotcms>/WEB-INF/velocity

for jar in <plugins>/*.jar; do
  echo "--- $jar"
  unzip -p "$jar" '*.class' 2>/dev/null | strings \
    | grep -o "esSearch\|esSearchRaw\|getEsSearchAPI\|IndiciesAPI\|ESSearchResults\|createSnapshot\|uploadSnapshot" | sort -u
done

Read the startup validation result

docker compose logs dotcms | grep -i "IndexStartupValidator"
docker compose logs dotcms | grep -i -E "OSIndexAPIImpl|SystemExitManager"

(Adapt to however you read logs. The class names are what matter.)


Startup Validation#


dotCMS validates its connection to OpenSearch every time it starts. Understanding this is the difference between a migration you can trust and one you are guessing at.

What It Checks#

Three things, logged individually:

INFO IndexStartupValidator - OS version check passed: 3.8.0
INFO IndexStartupValidator - Endpoint separation check passed. ES: [...] — OS: [...]
INFO IndexStartupValidator - OpenSearch startup validation PASSED — connected to OS successfully;
     migration phase PHASE_2_DUAL_WRITE_OS_READS is active.

That the target is OpenSearch 3.x. That the two engines are not the same endpoint. That dotCMS can actually connect and authenticate.

The third line names the active position in plain words. Learn what a healthy boot looks like early, so that its absence registers.

On Failure, by Position#

At positions 1 and 2, a validation failure is not fatal. dotCMS logs it at ERROR, resets the position to 0, and continues serving from Elasticsearch:

ERROR IndexStartupValidator - OpenSearch startup validation FAILED — halting OS migration;
      dotCMS falls back to ES-only (PHASE_0_MIGRATION_NOT_STARTED): <reason>
WARN  IndexConfigHelper$MigrationPhase - Migration phase reset to PHASE_0_MIGRATION_NOT_STARTED
      (was PHASE_1_DUAL_WRITE_ES_READS).

This is reasonable — Elasticsearch is current at those positions, so falling back to it is safe. But note what it means for you: the migration switches itself off and nothing user-visible changes. No readiness blocker, no banner, no failed page. The environment variable still says what you set. The only two ways to know are that log line and the current field in the readiness report.

That is the entire reason 3.3 and 4.3 exist as their own sub-steps. 4.3 is the one that matters most, because at position 2 a silent demotion produces a site that behaves perfectly — Elasticsearch is still complete — so you can test extensively and learn nothing.

At position 3 it is fatal, deliberately. OpenSearch is the primary store and Elasticsearch is no longer maintained, so falling back would mean serving stale content and writing new content into a decommissioned engine. Instead dotCMS retries and then stops:

FATAL OSIndexAPIImpl - OpenSearch is not reachable after 24 attempt(s).
      phase=PHASE_3_OPENSEARCH_ONLY, endpoints=[...],
      likelyCause=AUTH_FORBIDDEN (authentication/authorization rejected (HTTP 401/403) —
      verify the OS credentials and that the user has the required cluster permissions.),
      error=… — OS is the primary store in PHASE_3_OPENSEARCH_ONLY; cannot fall back to ES.
      Fix the connection per the likely cause above, then restart dotCMS.
INFO  SystemExitManager - System exit requested: … (exit code: 1)

The message names a likelyCause and what to do about it. Fix that and dotCMS starts.

What This Means After#

From position 3 onward, OpenSearch is a hard startup dependency. Not "search degrades" — dotCMS does not start. Under an orchestrator that is a crash loop rather than a slow site.

Three practical consequences:

  • Rotating the OpenSearch service account credential can take the CMS down. Treat that password with the same care as your database password, because it now occupies the same position.
  • An OpenSearch maintenance window is a dotCMS maintenance window if dotCMS might restart during it.
  • Your search cluster is production infrastructure. Monitoring, backups, disk headroom and upgrade planning for OpenSearch belong in the same tier as your database from Step 5 onward.

The Shadow-Write Warning#

Separate from startup, and the one runtime log line worth watching for throughout positions 1 and 2:

WARN ContentletIndexAPIImpl - OS shadow write failed in putToIndex —
     OS index may diverge until next reindex. Cause: <host>

Treat this as an action item. A document reached one engine and not the other. The shadow write is deliberately fire-and-forget so that a problem with the secondary engine cannot break a working site — which also means it is not retried and not queued for recovery. The gap is permanent until a full reindex.

If you see this line: note when it happened, find out what was wrong with the other engine at that moment, fix it, and run a full reindex before advancing a position.

It appears at WARN, which is on by default. If you do not see it and expect to, dotmarketing-config.properties documents the level under DOTCMS_SHADOW_WRITE_LOG_LEVEL — set as an environment variable it takes the DOT_ prefix.

Both points are settled in source: the code reads Config.getStringProperty("DOTCMS_SHADOW_WRITE_LOG_LEVEL", "WARN"), and its level switch falls through to Logger.warn for any unrecognised value — so the warning is at WARN under every configuration, including a misspelled one.