🎥 The YouTube version is also available:
Semantic Chunk for Manufacturing: Build Trustworthy Knowledge Without Exposing Data

When manufacturers consider introducing Semantic Chunks or Knowledge Graphs, they quickly encounter a major concern. Maintenance records, quality incident reports, MES and ERP data, BOMs, engineering changes, drawings, and process conditions often contain customer-specific specifications and proprietary manufacturing know-how. It is not always acceptable to send such information to an external LLM.
The important point is not to equate Semantic Chunks with “letting an LLM read documents.” A Semantic Chunk is, fundamentally, a way of organizing information into meaningful units—such as equipment, parts, processes, events, actions, and results—while preserving evidence that allows the original source to be verified later. An LLM can be a powerful component for generating candidates, but it is not a prerequisite.
A practical design separates public information from internal information. Public documents can be used to build general knowledge. Confidential internal data should remain in a closed environment and be handled with rules, dictionaries, BERT-based models, statistical methods, Knowledge Graphs, GNNs, and related techniques. The two sides should be connected not by transferring source documents, but through approved shared concepts.
Do Not Merge the Two Graphs
First, divide the data into two layers.
| Layer | Primary Data | Primary Processing | External Transmission |
|---|---|---|---|
| Public Knowledge Layer | Public standards, research papers, patents, government materials, public manuals | Generate candidates for concepts and relationships using LLMs and other methods; human review and adoption | Permitted, subject to terms of use |
| Internal Evidence Layer | Maintenance records, quality records, MES, ERP, BOM, QMS, engineering changes | Rules, dictionaries, BERT, classifiers, statistics, GNNs, internal search | Not permitted |
| Semantic Bridge | Shared terminology, classifications, concept IDs, mapping rules | Concept mapping and approval workflows | Source documents are never transferred |
On the public side, for example, general knowledge related to the concept of a “bearing” can be organized around failure modes, observed indicators, and possible actions.
Bearing
├─ Failure modes: wear, insufficient lubrication, overheating, vibration
├─ Observed indicators: temperature, vibration, noise
├─ Candidate actions: lubrication, replacement, alignment, load reduction
└─ Related equipment: motors, spindles, rotating machinery
The internal side, in contrast, retains the actual equipment, production lines, parts, measured values, and customer specifications.
Equipment_EQ001
──uses part──> Part_P-782
──has event──> Anomaly_EVT-456
──action taken──> Maintenance_Task_M-91
──result──> Downtime, recurrence status, quality impact
Even when internal parts or anomalies are linked to public concepts, there is no need to send actual names or numerical values to the public side. The internal mapping table only needs to record relationships such as the following.
Part_P-782 ──broaderMatch──> Public concept “Bearing”
Anomaly_EVT-456 ──observed-as──> Public concept “Overheating”
With this structure, users can refer to general knowledge such as “which indicators should be checked when overheating occurs” or “which actions may be relevant.” At the same time, information about which plant, which equipment, and which values were observed never leaves the organization.
Build the Source-and-Evidence Foundation Before the LLM Foundation
In the first stage of implementation, what should be fixed is not the model, but the source, permissions, and evidence. A Semantic Chunk should not store only a summary. It must always preserve where in the original source the information came from.
Chunk ID: SC-2026-00124
Event: Spindle temperature anomaly
Target: Equipment_EQ001
Temporary action: Changed operating conditions; replaced part
Confidence: 0.82
Evidence: Maintenance Record MR-2026-054, page 3, paragraph 2, table 4
Access permissions: Production Engineering and Maintenance Departments
If the original source, version, document ID, permissions, evidence location, and extraction, correction, and approval history are retained internally, reproducibility is not lost even if the model changes or an extraction result is corrected. Conversely, a design that preserves only AI-generated summaries cannot become a knowledge infrastructure that supports quality assurance, audits, or recurrence prevention.
The Core of Building Semantic Chunks Without an LLM
Manufacturing forms and reports already contain substantial structure: equipment numbers, part numbers, anomaly codes, dates, processes, responsible personnel, and action categories. These can first be extracted through existing master data, form definitions, regular expressions, dictionaries, OCR, and layout analysis.
| Purpose | Methods Used Without an LLM |
|---|---|
| Splitting forms and tables | Form definitions, heading rules, OCR, layout analysis |
| Extracting equipment, parts, and processes | Master-data matching, regular expressions, dictionaries, named-entity-recognition models |
| Classifying free text | Japanese BERT-based classifiers, similarity search |
| Searching for similar failures | Embedding models, vector search, BM25 |
| Anomaly detection and early warning | Time-series statistics, Isolation Forest, autoencoders |
| Estimating possible causes | Rules, decision trees, Bayesian networks, causal analysis |
| Exploring relationships and impact scope | Knowledge Graphs, GNNs, link prediction |
If a maintenance record states, “The bearing was replaced because vibration was high,” there is no need to automatically determine the cause at the initial stage. It is sufficient to separate confirmed and unconfirmed information as follows.
Target equipment: Confirmed from equipment number
Candidate event: Vibration anomaly
Candidate action: Bearing replacement
Source text: Preserved as-is
Evidence location: Free-text field
Unconfirmed items: Root cause, confirmation of effect
Keeping unconfirmed information explicitly unconfirmed prevents the creation of Knowledge Graphs based on unsupported assumptions. AI and rules may generate candidates, but they must not register unsupported causes or causal relationships as facts.
The Human Role Is Not Data Entry for Every Record, but Judgment on Standards and Exceptions
It is not realistic to construct every public Knowledge Graph and internal mapping table manually. The first thing people should define is the minimum shared vocabulary required for the target operation. For maintenance work, the following scope is sufficient at the beginning.
Equipment types: Pumps, motors, spindles, compressors
Part types: Bearings, belts, seals, gears
Events: Overheating, vibration, abnormal noise, leakage, stoppage
Actions: Inspection, lubrication, adjustment, replacement, shutdown
Results: Restored, observation required, recurrence, quality impact
Using this shared vocabulary as a standard, machines generate candidates and people approve only the uncertain or high-impact cases.
| Stage | What the Machine Does | What People Do |
|---|---|---|
| Master-data use | Retrieves part codes, item names, and classification codes from existing master data | Verifies master-data quality |
| Rule-based candidates | Generates candidates from terms and classification codes | Approves initial rules |
| Similarity-based candidates | Presents top candidates using BERT embeddings or item-name similarity | Selects only ambiguous candidates |
| Confidence routing | Sorts results into high, medium, and low confidence | Reviews medium- and low-confidence items |
| Reuse | Feeds approved results back into dictionaries and classifiers | Oversees new and exceptional cases |
For example, “Deep-groove ball bearing 6205” can be linked to “bearing” with high confidence. However, a “rotary support module Z7” or a customer-specific part must not be definitively categorized based on an inferred meaning. To represent this difference, relationships should not be limited to a simple is-a structure.
| Relationship | Meaning | Treatment |
|---|---|---|
exactMatch |
The same concept | Easier to automate when confidence is high |
broaderMatch |
Mapping to a broader concept | Used to reference general knowledge |
relatedMatch |
A related candidate | Not used alone as the basis for recommendations |
unmapped |
No mapping or pending decision | Not connected; sent for review |
Human review should focus on new part types, processes, or defect categories; mappings that affect safety, quality assurance, or customer reporting; low-confidence candidates; and candidates that conflict with established knowledge. This allows expert time to be concentrated not on entering every record, but on the areas where meaning and responsibility matter most.
Use GNNs After an Evidence-Grounded Graph Is in Place
GNNs can help infer similarities and undiscovered relationships among equipment, parts, processes, defects, maintenance activities, and quality results. However, simply building a graph does not automatically create correct knowledge. Definitions of nodes and edges, time context, permissions, and source evidence are required first.
Once an evidence-grounded graph is in place, GNNs and statistical models can present candidates such as:
- Anomalies likely to occur in similar equipment
- Potential effects of a part change on downtime or quality
- Unregistered equipment-to-part or event-to-action relationships
- Combinations of processes and conditions associated with recurrence
However, these outputs must be treated as candidates for verification, not as instructions for automatically executing corrective actions. Users should be shown not only predictions, but also the records referenced, graph paths, and source evidence.
A Phased Implementation Sequence
Do not begin by targeting all company-wide drawings, BOMs, quality records, and maintenance records. Start with a single operation: maintenance records from one plant, for example, or quality incident reports for one product family. Limit the period to the most recent one to three years and begin with several thousand to several tens of thousands of records.
The evaluation should not focus only on AI accuracy.
- How much was the time required to find past cases or prepare reports reduced?
- How often did people need to correct the extraction of equipment, anomalies, and actions?
- Could the system correctly present source evidence?
- Did it prevent the display of documents outside the user’s permissions or information across customers?
- Which free-text cases genuinely require a closed-network LLM in the future?
Only after reviewing these results can an organization determine where it is worth investing in a closed-network LLM or a dedicated cloud environment. A closed-network LLM should not be an assumed prerequisite. It should be added only when it is confirmed that free-text content cannot be handled adequately with rules or BERT-based methods, and must be processed without leaving the organization.
Conclusion
The first question in manufacturing Semantic Chunk initiatives should not be, “Which LLM should we use?” It should be: which meanings should be shared, and which judgments should remain with people, while preserving source documents, permissions, evidence, and approval history inside the organization?
Public information can be used to organize general concepts and reference evidence. Internal information can be handled in a closed environment using structured data, rules, BERT, machine learning, Knowledge Graphs, and GNNs. By connecting the two through shared concept IDs, organizations can use general knowledge and their own operational experience within the same analytical and decision-making framework—without exposing confidential information externally.
An LLM does not replace this foundation. Positioning it as a replaceable component for candidate generation, used only where needed, is what creates a knowledge infrastructure that remains safe, sustainable, and operational over the long term.

Chinoba
Intelligence as Relationship
Research Platform
founded by
Masao Watanabe
AI Systems Architecture
Decision Trace
Human–AI Coordination
Algorithmic Governance
Related Research
This topic is part of the Chinoba Knowledge Base.

コメント