Building Semantic Chunks and Knowledge Graphs in Manufacturing Without Exposing Confidential Data

Knowledge Base Archive This article is part of the Chinoba Knowledge Base. Explore Chinoba.org →

🎥 The YouTube version is also available:

Semantic Chunk for Manufacturing: Build Trustworthy Knowledge Without Exposing Data

Semantic Chunk for Manufacturing: Build Trustworthy Knowledge Without Exposing Data

Books: Semantic Chunk Building from Evidence-Grounded Semantic Units: Transforming documents, tables, and business data into explainable, executable knowledge

When manufacturers consider introducing Semantic Chunks or Knowledge Graphs, they quickly encounter a major concern. Maintenance records, quality incident reports, MES and ERP data, BOMs, engineering changes, drawings, and process conditions often contain customer-specific specifications and proprietary manufacturing know-how. It is not always acceptable to send such information to an external LLM.

The important point is not to equate Semantic Chunks with “letting an LLM read documents.” A Semantic Chunk is, fundamentally, a way of organizing information into meaningful units—such as equipment, parts, processes, events, actions, and results—while preserving evidence that allows the original source to be verified later. An LLM can be a powerful component for generating candidates, but it is not a prerequisite.

A practical design separates public information from internal information. Public documents can be used to build general knowledge. Confidential internal data should remain in a closed environment and be handled with rules, dictionaries, BERT-based models, statistical methods, Knowledge Graphs, GNNs, and related techniques. The two sides should be connected not by transferring source documents, but through approved shared concepts.

Do Not Merge the Two Graphs

First, divide the data into two layers.

Layer Primary Data Primary Processing External Transmission
Public Knowledge Layer Public standards, research papers, patents, government materials, public manuals Generate candidates for concepts and relationships using LLMs and other methods; human review and adoption Permitted, subject to terms of use
Internal Evidence Layer Maintenance records, quality records, MES, ERP, BOM, QMS, engineering changes Rules, dictionaries, BERT, classifiers, statistics, GNNs, internal search Not permitted
Semantic Bridge Shared terminology, classifications, concept IDs, mapping rules Concept mapping and approval workflows Source documents are never transferred

On the public side, for example, general knowledge related to the concept of a “bearing” can be organized around failure modes, observed indicators, and possible actions.

Bearing
  ├─ Failure modes: wear, insufficient lubrication, overheating, vibration
  ├─ Observed indicators: temperature, vibration, noise
  ├─ Candidate actions: lubrication, replacement, alignment, load reduction
  └─ Related equipment: motors, spindles, rotating machinery

The internal side, in contrast, retains the actual equipment, production lines, parts, measured values, and customer specifications.

Equipment_EQ001
  ──uses part──> Part_P-782
  ──has event──> Anomaly_EVT-456
  ──action taken──> Maintenance_Task_M-91
  ──result──> Downtime, recurrence status, quality impact

Even when internal parts or anomalies are linked to public concepts, there is no need to send actual names or numerical values to the public side. The internal mapping table only needs to record relationships such as the following.

Part_P-782 ──broaderMatch──> Public concept “Bearing”
Anomaly_EVT-456 ──observed-as──> Public concept “Overheating”

With this structure, users can refer to general knowledge such as “which indicators should be checked when overheating occurs” or “which actions may be relevant.” At the same time, information about which plant, which equipment, and which values were observed never leaves the organization.

Build the Source-and-Evidence Foundation Before the LLM Foundation

In the first stage of implementation, what should be fixed is not the model, but the source, permissions, and evidence. A Semantic Chunk should not store only a summary. It must always preserve where in the original source the information came from.

Chunk ID: SC-2026-00124
Event: Spindle temperature anomaly
Target: Equipment_EQ001
Temporary action: Changed operating conditions; replaced part
Confidence: 0.82
Evidence: Maintenance Record MR-2026-054, page 3, paragraph 2, table 4
Access permissions: Production Engineering and Maintenance Departments

If the original source, version, document ID, permissions, evidence location, and extraction, correction, and approval history are retained internally, reproducibility is not lost even if the model changes or an extraction result is corrected. Conversely, a design that preserves only AI-generated summaries cannot become a knowledge infrastructure that supports quality assurance, audits, or recurrence prevention.

The Core of Building Semantic Chunks Without an LLM

Manufacturing forms and reports already contain substantial structure: equipment numbers, part numbers, anomaly codes, dates, processes, responsible personnel, and action categories. These can first be extracted through existing master data, form definitions, regular expressions, dictionaries, OCR, and layout analysis.

Purpose Methods Used Without an LLM
Splitting forms and tables Form definitions, heading rules, OCR, layout analysis
Extracting equipment, parts, and processes Master-data matching, regular expressions, dictionaries, named-entity-recognition models
Classifying free text Japanese BERT-based classifiers, similarity search
Searching for similar failures Embedding models, vector search, BM25
Anomaly detection and early warning Time-series statistics, Isolation Forest, autoencoders
Estimating possible causes Rules, decision trees, Bayesian networks, causal analysis
Exploring relationships and impact scope Knowledge Graphs, GNNs, link prediction

If a maintenance record states, “The bearing was replaced because vibration was high,” there is no need to automatically determine the cause at the initial stage. It is sufficient to separate confirmed and unconfirmed information as follows.

Target equipment: Confirmed from equipment number
Candidate event: Vibration anomaly
Candidate action: Bearing replacement
Source text: Preserved as-is
Evidence location: Free-text field
Unconfirmed items: Root cause, confirmation of effect

Keeping unconfirmed information explicitly unconfirmed prevents the creation of Knowledge Graphs based on unsupported assumptions. AI and rules may generate candidates, but they must not register unsupported causes or causal relationships as facts.

The Human Role Is Not Data Entry for Every Record, but Judgment on Standards and Exceptions

It is not realistic to construct every public Knowledge Graph and internal mapping table manually. The first thing people should define is the minimum shared vocabulary required for the target operation. For maintenance work, the following scope is sufficient at the beginning.

Equipment types: Pumps, motors, spindles, compressors
Part types: Bearings, belts, seals, gears
Events: Overheating, vibration, abnormal noise, leakage, stoppage
Actions: Inspection, lubrication, adjustment, replacement, shutdown
Results: Restored, observation required, recurrence, quality impact

Using this shared vocabulary as a standard, machines generate candidates and people approve only the uncertain or high-impact cases.

Stage What the Machine Does What People Do
Master-data use Retrieves part codes, item names, and classification codes from existing master data Verifies master-data quality
Rule-based candidates Generates candidates from terms and classification codes Approves initial rules
Similarity-based candidates Presents top candidates using BERT embeddings or item-name similarity Selects only ambiguous candidates
Confidence routing Sorts results into high, medium, and low confidence Reviews medium- and low-confidence items
Reuse Feeds approved results back into dictionaries and classifiers Oversees new and exceptional cases

For example, “Deep-groove ball bearing 6205” can be linked to “bearing” with high confidence. However, a “rotary support module Z7” or a customer-specific part must not be definitively categorized based on an inferred meaning. To represent this difference, relationships should not be limited to a simple is-a structure.

Relationship Meaning Treatment
exactMatch The same concept Easier to automate when confidence is high
broaderMatch Mapping to a broader concept Used to reference general knowledge
relatedMatch A related candidate Not used alone as the basis for recommendations
unmapped No mapping or pending decision Not connected; sent for review

Human review should focus on new part types, processes, or defect categories; mappings that affect safety, quality assurance, or customer reporting; low-confidence candidates; and candidates that conflict with established knowledge. This allows expert time to be concentrated not on entering every record, but on the areas where meaning and responsibility matter most.

Use GNNs After an Evidence-Grounded Graph Is in Place

GNNs can help infer similarities and undiscovered relationships among equipment, parts, processes, defects, maintenance activities, and quality results. However, simply building a graph does not automatically create correct knowledge. Definitions of nodes and edges, time context, permissions, and source evidence are required first.

Once an evidence-grounded graph is in place, GNNs and statistical models can present candidates such as:

  • Anomalies likely to occur in similar equipment
  • Potential effects of a part change on downtime or quality
  • Unregistered equipment-to-part or event-to-action relationships
  • Combinations of processes and conditions associated with recurrence

However, these outputs must be treated as candidates for verification, not as instructions for automatically executing corrective actions. Users should be shown not only predictions, but also the records referenced, graph paths, and source evidence.

A Phased Implementation Sequence

Do not begin by targeting all company-wide drawings, BOMs, quality records, and maintenance records. Start with a single operation: maintenance records from one plant, for example, or quality incident reports for one product family. Limit the period to the most recent one to three years and begin with several thousand to several tens of thousands of records.

The evaluation should not focus only on AI accuracy.

  • How much was the time required to find past cases or prepare reports reduced?
  • How often did people need to correct the extraction of equipment, anomalies, and actions?
  • Could the system correctly present source evidence?
  • Did it prevent the display of documents outside the user’s permissions or information across customers?
  • Which free-text cases genuinely require a closed-network LLM in the future?

Only after reviewing these results can an organization determine where it is worth investing in a closed-network LLM or a dedicated cloud environment. A closed-network LLM should not be an assumed prerequisite. It should be added only when it is confirmed that free-text content cannot be handled adequately with rules or BERT-based methods, and must be processed without leaving the organization.

Conclusion

The first question in manufacturing Semantic Chunk initiatives should not be, “Which LLM should we use?” It should be: which meanings should be shared, and which judgments should remain with people, while preserving source documents, permissions, evidence, and approval history inside the organization?

Public information can be used to organize general concepts and reference evidence. Internal information can be handled in a closed environment using structured data, rules, BERT, machine learning, Knowledge Graphs, and GNNs. By connecting the two through shared concept IDs, organizations can use general knowledge and their own operational experience within the same analytical and decision-making framework—without exposing confidential information externally.

An LLM does not replace this foundation. Positioning it as a replaceable component for candidate generation, used only where needed, is what creates a knowledge infrastructure that remains safe, sustainable, and operational over the long term.

Related Research

This topic is part of the Chinoba Knowledge Base.

Chinoba Research
Chinoba-lab Open Source
Books and Library

コメント

タイトルとURLをコピーしました