Issue
After upgrading Fusion or making changes to the SharePoint Optimized Connector, some documents—especially non-HTML types (PDF, Office)—are missing expected content fields such as body_t or custom property fields in the Solr index. In some scenarios, fields are present in one version of the document and missing in another, or documents are indexed with partial content only. This commonly occurs when attempting to use asynchronous parsing (async parsing) with the SharePoint connector in Fusion 5.8 through 5.11.
Diagnosis
Non-HTML documents are missing the
body_tfield or other custom fields after ingestion with the SharePoint connector.When using async parsing, documents may appear twice in the index—once with
Field.*properties but nobody_t, and once withbody_tbut missing custom fields.-
Errors may appear in logs, such as:
500 Internal Server Error: ... /rmeta/body ...Validation errors for fields with unsupported characters.
Switching between async and non-async parsing, or changing pipeline stages, produces inconsistent results.
The
error_message_tfield may be populated, indicating a downstream issue.
Environment
Lucidworks Fusion 5.8.x–5.11.x (Self-hosted, Kubernetes/EKS)
SharePoint Optimized Connector (version 2.1.0)
Typically seen in environments using document pipelines with custom parsing requirements
Cause
Async parsing limitations: In Fusion 5.8 through 5.11, async parsing supports only the Tika parser. Parser configuration in the data source is ignored when async parsing is enabled. Other parsers, including HTML and JSON, are not supported in async parsing until Fusion 5.12 and later.
Pipeline configuration: Async parsing requires the Solr Partial Update Indexer stage, with specific options enabled. Using only the Solr Indexer or misconfiguring pipeline options will result in partial or missing fields.
Deprecated parser stages: The "Apache Tika" and "Apache Tika Server" stages are deprecated and may not be maintained for future compatibility. Their behavior can differ from the _system parser.
Document type and field expectations: Some SharePoint libraries are not document content sources and will not have a
body_tfield by design.Connector patch compatibility: Applying a patch built for an older Fusion version to a newer environment can cause unexpected connector or pipeline issues.
Resolution
Verify Fusion and connector version compatibility
Ensure that any patches applied to the connectors-backend or connector-plugin images are built specifically for your Fusion version.
Do not use a patch intended for an earlier Fusion release.
Understand async parsing limitations
In Fusion 5.8 through 5.11, only the Tika parser is supported for async parsing. All other parser stages are ignored when async parsing is enabled.
Support for HTML, JSON, and other parsers in async parsing is available starting in Fusion 5.12.
Configure the index pipeline for async parsing
Open the Index Pipelines interface in Fusion.
-
Ensure the pipeline used by the SharePoint connector has the following configuration:
-
Stages to include (in order):
Field Mapping (as needed)
Solr Dynamic Field Name Mapping (if using dynamic fields)
Solr Partial Update Indexer (this must be the final stage)
-
-
In the Solr Partial Update Indexer stage:
Disable Reject Update if Solr Document is not Present
Enable Process All Pipeline Doc Fields
Enable Allow reserved fields
Add at least one update field (for example, an increment field like
docs_counter_iwith value 1)Remove or disable the regular Solr Indexer stage
Save the pipeline and ensure the SharePoint data source is configured to use async parsing, if needed.
Example settings for the Solr Partial Update Indexer stage:
Identify content documents correctly
Use the
_lw_document_type_sfield in Solr to filter for actual content documents (such as those withbody_tfields), and distinguish from library or metadata-only documents.
Troubleshoot missing fields
-
If documents still lack
body_tor custom fields:Review pipeline logs for errors such as schema validation issues or HTTP 500 errors.
Ensure all required pipeline stages are present and correctly ordered.
Confirm that the SharePoint connector is connecting and authenticating successfully (check connector-plugin and connectors-backend logs).
Check for the presence of the
error_message_tfield, which can indicate a server-side or SharePoint-specific error. Review server logs for details.
Workarounds for Fusion 5.8–5.11
For document types that require parsing beyond Tika (e.g., HTML or JSON), use non-async parsing until upgrading to Fusion 5.12 or later.
Consider restructuring parser stages to avoid deprecated Tika stages where possible.
Additional notes
Note: "Apache Tika" and "Apache Tika Server" parser stages are marked deprecated and may be removed in future Fusion releases. Migrate parsing requirements to supported parsers as new releases become available.
Note: Some SharePoint document libraries are not intended to contain document bodies. The absence of body_t may be expected for these documents.
Guide to upgrading for full async parsing support
If your workflow requires async parsing for HTML, JSON, or other parser types, plan to upgrade Fusion to version 5.12 or above.
Review release notes and compatibility guidance before upgrading, and ensure all pipeline and connector plugins are updated accordingly.
Additional resources