🎥 The YouTube version is also available:
Semantic Chunk for Manufacturing: Build Trustworthy Knowledge Without Exposing Data

When using generative AI in manufacturing operations, one question inevitably arises:
How much of our design documentation, maintenance records, VOC data, and customer specifications can we provide to AI?
The key is not to begin with a design that lets AI “read all company documents.” Original documents should remain under the organization’s control, while only the minimum necessary evidence is temporarily provided to the LLM according to the user’s permissions and purpose of use.
The unit that implements this boundary is the Semantic Chunk.
A Semantic Chunk is not merely a method for splitting content to improve RAG retrieval accuracy. It is also a boundary for knowledge and security—one that controls what information AI may reference and makes it possible to verify the evidence afterward.
Simply Splitting Content into Smaller Pieces Does Not Make It Safe
If a long PDF is divided into chapters or paragraphs and only small chunks are passed to an LLM, the exposure surface is smaller than sending the entire document. This is important.
However, segmentation alone is not security.
If a chunk contains a customer name, drawing number, equipment configuration, export-control information, or personal data, that chunk is still sensitive. Likewise, if a document the user is not authorized to access is included in the search results, it is a data leak—regardless of how accurate the retrieval may be.
The following order is essential:
- Before retrieval, narrow the scope to documents the user is authorized to read.
- Before generation, minimize the information sent to the LLM.
- After the response, make the evidence and processing history available for review.
Semantic Chunks become the smallest unit that connects these three requirements.
Send Authorized Evidence, Not Entire Documents
Suppose a user asks:
“Has a similar vibration anomaly occurred in Equipment X in the past?”
An undesirable design would send the full maintenance-report PDF—or an entire folder—to the LLM. It could include customer information, worker information, and records for other equipment that have nothing to do with the question.
Instead, the system should operate as follows:
- Verify the user’s department, role, project, site, and viewing permissions.
- Restr retrieval to materials the user is authorized to access.
- Retrieve only the paragraphs relevant to “Equipment X,” “vibration anomaly,” “cause,” and “countermeasure.”
- Mask personal names and contact details where necessary.
- Send only the question and the authorized evidence chunks to the LLM.
- Attach the document ID, revision, page, and relevant location to the response.
The LLM does not have unrestricted access to the document repository. It reads only the material provided at that point in time and responds within the available evidence.
Information That a Semantic Chunk Should Contain
A chunk is not merely a fragment of text. At a minimum, it should include the following metadata together with its content.
| Item | Example |
|---|---|
| Source identifier | MNT-2026-018 / Rev.3 / p.12 |
| Responsible organization / project | XX Business Division, Project A |
| Confidentiality classification | Public, Internal, Confidential, Highly Confidential |
| Regulatory classification | General, Export-Controlled, Defense-Related, Nuclear Security-Related |
| AI-use eligibility | Allowed, Conditional, Prohibited |
| Access conditions | Maintenance team, Design owner, Project participants |
| Source location | Heading, page, paragraph, table number |
| Evidence link | Link to the original document in the document-management system |
The important point is not to retrieve content merely because it is semantically similar. The system must first decide that the content may be retrieved, and only then include it in the searchable scope.
Vector search and keyword search should run only after filtering through an ACL (Access Control List). Even if the search engine finds highly relevant content, the chunk must not appear in the results if the user is not authorized to view it.
Highly Confidential Information Must Be Blocked by the System, Not by Warnings
For highly confidential materials, simply warning users is insufficient. The system must reject them at every stage: document registration, labeling, storage location, IAM permissions, ingestion processing, retrieval, and LLM invocation.
At document-registration time, the following should be mandatory attributes.
| Item | Example |
|---|---|
| Confidentiality classification | Public / Internal / Confidential / Highly Confidential |
| Regulatory classification | General / Export-Controlled / Defense-Related / Nuclear Security-Related |
| Customer restriction | None / AI-use restricted / Individual approval required |
| AI-use eligibility | Allowed / Conditional / Prohibited |
| Approval information | Approver, approval date, expiration date |
For an initial implementation, the policy can be kept simple:
- Materials classified as
Highly Confidential,Export-Controlled,Defense-Related,Nuclear Security-Related, orAI Use Prohibitedmust not be registered in RAG systems or sent to an LLM. - Materials classified as
ConfidentialorIndividual Approval Requiredshould be prohibited by default, with exceptions allowed only for approved projects, purposes, and periods. - Materials classified as
Internalor below may be used, subject to ACL verification, masking, and minimum necessary disclosure.
It is not enough to hide prohibited materials from the search screen. They should not be chunked, vectorized, or indexed at all. Keeping them out of the retrieval database greatly reduces the risk of a later path that accidentally sends them to an LLM.
Protect Information Through “Isolated Shelves” and Least Privilege
Do not make all existing documents available to AI search from the outset. Divide storage into three areas:
ai-eligible: Materials approved for AI retrieval and LLM useai-review: Materials awaiting classification, contract review, or access verificationai-excluded: Highly confidential, regulated, or AI-prohibited materials
The RAG ingestion role should be granted permission to read only ai-eligible. It should have no access to ai-excluded, not only through application-level conditional logic but also at the IAM level.
With this design, even if there is an application defect, the ingestion process cannot read excluded material. Security should not depend on a single decision rule; it should be protected across multiple layers.
Manage Exceptional Use Through Time-Limited “Exception Tokens”
There may be cases where confidential materials must be used. If exceptions are granted only through verbal communication or email, it becomes unclear who may use which materials, for what purpose, and for how long.
Instead, the results of an individual review should be retained as an approval record within the system.
| Approval item | Example |
|---|---|
| Scope | Document ID, project ID, chunk range |
| Purpose of use | Maintenance-history search assistance only |
| Users | Specified team and roles |
| Permitted model | Claude Haiku on Amazon Bedrock only |
| Transmission limit | Maximum five chunks; masking required |
| Valid period | 2026-10-01 to 2026-12-31 |
| Approvers | Information-management owner and export-control representative |
At query time, the system adds the material to the searchable scope only when the target, user, purpose, and validity period match the approval record. When the period expires, access stops automatically.
This is not a model in which “once approved, always available.” It creates the smallest possible, revocable permission bound to a specific purpose.
Block Again Immediately Before LLM Transmission
The final boundary should be placed immediately before invoking the LLM.
Even if an inappropriate chunk is accidentally included in the retrieval results, the pre-transmission policy check should reject it.
if chunk.ai_use != "allowed":
block
if chunk.regulatory in prohibited_categories:
block
if approval is missing or expired:
block
if user role / project scope is invalid:
block
When access is rejected, the system should not reveal the document’s contents. It should provide guidance such as:
This material is outside the scope of AI-assisted use. If you need to use it, please request an individual review from the information-management owner.
The LLM should also receive fixed instructions: do not infer information that is not contained in the evidence; do not make decisions, approvals, or design changes; clearly state uncertainties; and always return the source document ID and relevant location.
Semantic Chunks Are the Smallest Unit of Knowledge—and the Smallest Unit of Accountability
The value of Semantic Chunks is not that they make answers shorter.
Their value lies in enabling the organization to trace which part of which original document and revision was retrieved; under whose authority; for what purpose; which model received it; and what response was returned.
This history is more than an access log. It becomes a Decision Trace that enables evidence verification, continuous improvement, and accountability.
The key to using generative AI safely is not to show AI everything. Control of original documents, permissions, and decision-making authority must remain with the organization. AI should assist with retrieval, summarization, and explanation only within the boundaries of authorized evidence.
Semantic Chunk is therefore both a RAG technology and a foundation for implementing knowledge governance in the AI era.

Chinoba
Intelligence as Relationship
Research Platform
founded by
Masao Watanabe
AI Systems Architecture
Decision Trace
Human–AI Coordination
Algorithmic Governance
Related Research
This topic is part of the Chinoba Knowledge Base.

コメント