Pergamum Pulse Ingest Workbench and Curator Operations Design
Version 0.1 — Governed Operational Design
This document defines the operational design for the Pergamum Pulse Ingest Workbench and the associated curator workflow.
It converts the Pergamum Pulse knowledge, rights, semantic, retrieval, change-intelligence and HECATE specifications into an executable human-and-machine operating model.
The governing operating principle is:
Machines SHALL perform deterministic, repetitive and evidence-extractable work. ZARA SHALL propose semantic interpretations and classifications. Curators SHALL review uncertainty, conflict, risk and authority-bearing Decisions.
Additional governing principles are:
The source artefact SHALL remain immutable.
Converted Markdown SHALL be treated as one transformation candidate, not as the source of record.
The canonical machine representation SHOULD be a governed document AST from which MDX, Source Fragments and Retrieval Chunks are projected.
Source Fragments SHALL follow stable source evidence boundaries. Retrieval Chunks SHALL remain purpose-specific derived objects.
ZARA proposals SHALL remain distinguishable from approved canonical records.
HECATE SHALL validate objective conformance and SHALL NOT absorb legal, semantic, governance, Publication or Runtime authority.
Kanban SHALL manage cases and responsibilities. Detailed source and document handling SHALL occur in a Case Workbench.
Every proposal, review action, exception, approval, rework action and canonical write SHALL preserve actor, evidence, time, reason and provenance.
Workflow SHALL be Profile-driven and reusable across regulatory, scientific, dataset, IoT, certification, internal, tenant and white-label knowledge.
1. Purpose
The Ingest Workbench SHALL convert captured source material into governed Pergamum Pulse knowledge without requiring curators to perform unnecessary clerical work.
The Workbench SHALL:
- capture source artefacts immutably;
- generate deterministic technical metadata;
- allow ZARA to propose evidence-backed semantic metadata;
- preserve extraction candidates and conflicts;
- support efficient review by exception;
- separate source structure from retrieval chunking;
- enforce rights, security, tenant and white-label boundaries;
- route authority-bearing Decisions to authorised Roles;
- validate objective completeness through HECATE;
- preserve full historical lineage;
- support replay, audit and revalidation;
- preserve both Module and Component lineage.
2. Scope
This design applies to:
- regulations;
- regulatory drafts;
- standards;
- guidance;
- methodologies;
- scientific literature;
- datasets;
- code lists;
- device manuals;
- product data sheets;
- API and protocol specifications;
- certificates;
- assurance evidence;
- supplier documentation;
- internal Viroway documents;
- canonical ZAYAZ documentation;
- tenant knowledge;
- white-label knowledge;
- public web content.
Supported source formats MAY include PDF, HTML, Markdown, MDX, DOCX, XLSX, CSV, JSON, XML, images, scans, API payloads and structured feeds.
This document does not itself grant rights, approve legal interpretations, approve semantic equivalence, issue Change Decisions, approve Publication or activate Runtime state.
3. Operating Model
Source Submission
↓
Immutable Capture
↓
Deterministic Technical Processing
↓
ZARA Proposal Generation
↓
Curator Verification by Exception
↓
Canonical Document AST
↓
Source Fragments
↓
HECATE Structural Admission
↓
Chunk Plan Proposal
↓
Retrieval Curator Review
↓
Semantic Extraction
↓
Domain Curator Review
↓
Impact and Change Intelligence
↓
Final Knowledge Admission
The operating model SHALL distinguish observation, extraction, proposal, validation, approval, admission, projection, release and activation.
Observed
≠ Extracted
≠ Proposed
≠ Validated
≠ Approved
≠ Admitted
≠ Projected
≠ Released
≠ Activated
4. Responsibility Model
4.1. Deterministic Services
Deterministic workers SHALL perform file hashing, malware scanning, MIME detection, file-size calculation, page counting, immutable storage, duplicate detection, identifier allocation, metadata extraction, native text extraction, layout extraction, bounding-box extraction, table extraction, OCR of image-only regions, AST construction, paragraph-number detection, heading hierarchy detection, offset calculation, digest generation, schema validation, database writes, event publication, lexical indexing, embedding generation, vector-index updates and HECATE execution.
Workers SHOULD be idempotent, retry-safe, versioned, observable, queue-driven and deterministic where technically feasible.
4.2. Specialised ZARA Agents
Specialised ZARA Agents MAY propose publisher, title, language, document type, publication status, publication date, Source Series, jurisdiction, authority classification, candidate rights holder, rights-risk classification, document structure, content classes, extraction reconciliation, conversion defects, chunk boundaries, Knowledge Assertions, Knowledge Entities, relationships, mappings, semantic differences, affected Modules, affected Components, Change Candidates and remediation suggestions.
Every ZARA proposal SHALL preserve proposed value, confidence, evidence, reasoning summary, Agent identity, Agent Profile, model, model version, prompt Profile, known limitations, review requirement and provenance.
4.3. Curators
Curators SHALL focus on uncertain identity, conflicting extraction, rights and permissions, semantic interpretation, applicability, equivalence, domain correctness, chunk exceptions, change impact, admission, exceptions and escalation.
Curators SHALL NOT be required to manually enter deterministic values already proven by the system.
4.4. HECATE
HECATE SHALL validate identifiers, required fields, object references, digests, lifecycle transitions, rights-state completeness, structural coverage, Source Fragment lineage, chunk lineage, tenant scope, white-label scope, Module lineage, Component lineage, workflow prerequisites, approval presence and receipt integrity.
HECATE SHALL NOT decide legal ownership, disputed rights, regulatory meaning, semantic equivalence, public Publication, certification or Runtime activation.
5. Automation Classes
| Class | Description | Example | Default Treatment |
|---|---|---|---|
D1 | Directly observable and deterministic | SHA-256 digest | Automatic |
D2 | Deterministic with governed lookup | Duplicate resolution | Automatic with receipt |
P1 | High-confidence evidence-derived proposal | Title from cover | Review by exception |
P2 | Semantic proposal requiring review | Authority class | Curator verification |
A1 | Authority-bearing Decision | Rights permission | Authorised human or workflow |
A2 | Constitutional or production-impacting Decision | Activation | Separate authority |
Automatic acceptance MAY be permitted for D1, D2 and selected P1 fields only under a governed automation policy defining confidence threshold, evidence requirements, source Profile, conflict absence, approved method, prior approved source pattern, fallback and audit treatment.
The interface SHOULD support grouped review by exception.
Every override SHALL preserve the original proposal, accepted value, curator, time, rationale, evidence, downstream impact and provenance.
6. Canonical Operational Objects
| Object | Prefix | Purpose |
|---|---|---|
| Ingest Case | PP-IC- | Governs one source-intake lifecycle |
| Ingest Profile | PP-ING- | Defines stages, Roles, requirements and Gates |
| Workflow Instance | PP-WFI- | Records one Profile execution |
| Workflow Stage | PP-WFS- | Records stage state |
| Workflow Task | PP-WFT- | Assigns required work |
| Proposal | PP-PROP- | Records a machine- or human-generated proposal |
| Proposal Evidence | PP-PEV- | Preserves evidence supporting a proposal |
| Review Decision | PP-RVD- | Records accept, modify, reject or escalate |
| Assignment | PP-ASG- | Records responsibility |
| Checklist Item | PP-CLI- | Records required operational evidence |
| Extraction Candidate | PP-EXT- | Records an extracted representation |
| Extraction Conflict | PP-ECF- | Records disagreement among candidates |
| Document Node | PP-DN- | Represents one AST node |
| Chunk Plan | PP-CHPL- | Groups proposed chunks for one purpose |
| Workbench Comment | PP-CMT- | Records contextual discussion |
| Stage Approval | PP-SAP- | Records authorised stage completion |
| Admission Decision | PP-ADM- | Records knowledge admission |
7. Ingest Case
ingest_case:
case_id: PP-IC-000001
case_version: 1.0.0
source_submission:
submitted_by: actor.example
intended_purpose:
- regulatory_intelligence
- change_monitoring
source_artifacts:
- PP-CAP-000001
ingest_profile:
profile_id: PP-ING-REGULATORY-DRAFT-001
profile_version: 1.0.0
current_stage: extraction_reconciliation
case_status: active
risk_level: high
tenant_scope:
mode: global_internal
provenance: {}
Every case SHALL identify stable identity, version, source submission, intended purpose, Ingest Profile, current stage, case status, risk, assignments, due dates, Findings, tenant scope, white-label scope and provenance.
Case states MAY include new, active, waiting, blocked, in review, returned for rework, approved, admitted, rejected, quarantined, withdrawn and historical.
8. Ingest Profile
Workflow SHALL be configuration-driven.
ingest_profile:
profile_id: PP-ING-REGULATORY-DRAFT-001
profile_version: 1.0.0
source_class: regulatory_standard_draft
stages:
- intake
- source_identity
- rights_review
- extraction_reconciliation
- structural_review
- chunk_review
- semantic_review
- change_intelligence
- final_admission
requirements:
official_source_verification: mandatory
rights_review: mandatory
paragraph_inventory: mandatory
application_requirement_inventory: conditional
table_reconciliation: mandatory
human_semantic_review: mandatory
production_activation: prohibited
roles:
source_identity: knowledge_curator
rights_review: rights_curator
extraction_reconciliation: document_curator
structural_review: knowledge_curator
chunk_review: retrieval_curator
semantic_review: domain_curator
change_intelligence: regulatory_architect
final_admission: knowledge_governance_approver
Initial Profiles SHOULD include regulatory draft, regulatory final, licensed standard, Viroway internal, IoT manual, dataset, scientific method, certificate and tenant document.
Material changes to stages, required fields, Roles, Gates, HECATE Profiles, approval policy or automation policy SHALL create a new Profile version.
9. Kanban Board
Default columns SHOULD include New Intake, Source Verification, Rights Review, Extraction Review, Structural Review, Chunk Review, Semantic Review, Impact Review, Final Admission, Completed and Quarantined.
Cards SHOULD show source title, publisher, source type, draft or final state, current stage, assigned Role, assigned actor, risk, due date, blocking Findings, proposal count, unresolved conflict count, HECATE state, completion percentage and tenant or white-label scope.
Drag-and-drop SHALL not bypass workflow Gates.
A card may advance only where required fields are complete, required approvals exist, blocking Findings are resolved or excepted, the HECATE Gate passes and the actor has permission.
10. Case Workbench
┌────────────────────┬────────────────────────────┬─────────────────────┐
│ Source Viewer │ Normalised Representation │ Stage Review Panel │
│ PDF / image / web │ AST / MDX / fragments │ Proposals / fields │
│ page navigation │ chunks / assertions │ evidence / actions │
└────────────────────┴────────────────────────────┴─────────────────────┘
The Source Viewer SHOULD support PDF rendering, page thumbnails, text selection, bounding-box highlighting, table highlighting, image-region highlighting, source locators, source version, source digest, citation copy and side-by-side comparison.
The centre pane SHOULD expose extraction candidates, canonical AST, reviewed MDX, Source Fragments, content classes, chunk plans, Knowledge Assertions, mappings, semantic differences and generated artefacts.
The review panel SHALL show required fields, ZARA proposals, evidence, confidence, HECATE Findings, reviewer guidance and actions to accept, modify, reject, escalate, comment, request reprocessing and approve stage.
11. Stage 1 — Intake
Required inputs SHALL include source artefact or URL, submitting actor, intended purpose, source class if known, tenant or global scope and confidentiality expectation.
The platform SHALL allocate an Ingest Case, capture the artefact, hash the file, scan for malware, detect media type, detect duplicate digest, record byte size, extract basic metadata, create a Source Capture candidate and emit intake events.
The Intake Gate SHALL require immutable capture, acceptable malware status, readable source, assigned case scope and no unresolved duplicate collision.
12. Stage 2 — Source Identity
ZARA MAY propose canonical title, publisher, language, publication date, publication status, Source Series, source type, jurisdiction, official source URL and version lineage.
The Knowledge Curator SHALL verify source identity, publisher, official status, draft or final state, Source Series relationship and duplicate or successor relationship.
The Source Identity Gate SHALL require approved identity or an explicit unresolved or quarantine Decision.
13. Stage 3 — Rights Review
The system MAY extract copyright notices, identify licence text, match Publisher Profiles, identify contractual references, classify restrictions, propose processing permissions and flag missing evidence.
The Rights Curator SHALL review the canonical Rights Record governed by PP-RLCPS and determine, verify or escalate the applicable:
- rights holder and copyright state;
- permission basis;
- rights scope;
processing_rights.storage;processing_rights.backup;processing_rights.parsing;processing_rights.ocr;processing_rights.translation;processing_rights.internal_summary;processing_rights.structured_extraction;processing_rights.semantic_indexing;processing_rights.embeddings;processing_rights.vector_storage;processing_rights.retrieval_augmented_generation;processing_rights.model_evaluation;processing_rights.fine_tuning;processing_rights.foundation_model_training;processing_rights.public_generation;processing_rights.derivative_dataset_creation;- applicable publication rights;
- attribution requirements;
- rights evidence;
- validity and expiry.
The Workbench SHALL NOT maintain independent aliases for canonical PP-RLCPS permission fields.
Unknown permission SHALL not be interpreted as granted.
The Rights Gate SHALL require sufficient disposition for operations needed by the next stage.
14. Stage 4 — Transformation and Extraction
The source artefact SHALL remain immutable.
The pipeline MAY use native PDF text extraction, layout-aware extraction, table extraction, image-region extraction, OCR for image-only regions, third-party PDF-to-Markdown conversion, HTML parsing and structured API extraction.
Multiple extraction candidates SHOULD be preserved.
Every Transformation SHALL identify input capture, input digest, output artefact, output digest, method, tool, version, provider, region, retention, rights eligibility, time and provenance.
The system SHOULD detect duplicated image text, malformed tables, broken headings, page headers and footers, broken ligatures, line-break hyphenation, missing sections, duplicate paragraphs, malformed Markdown, lost emphasis, broken references, OCR uncertainty and layout loss.
extraction_conflict:
conflict_id: PP-ECF-000001
source_locator:
pdf_page: 7
paragraph: 23
candidates:
native_pdf_text: "..."
markdown_converter: "..."
image_region_text: "..."
recommended_value: "..."
confidence: 0.982
review_required: true
The Document Curator SHALL resolve material conflicts and approve the canonical extraction representation.
15. Canonical Document AST
The canonical machine representation SHOULD be a governed document AST.
Node types MAY include document, cover, preamble, heading, paragraph, numbered paragraph, subparagraph, definition, Application Requirement, table, table row, table cell, list, list item, formula, note, footnote, appendix, figure, image, legal citation, cross-reference and page break.
document_node:
node_id: PP-DN-000001
node_version: 1.0.0
node_type: numbered_paragraph
semantic_role: normative
identifier:
source_label: "23"
content:
text: "Information is material when..."
structure:
heading_path:
- "3. Double materiality"
- "3.1. Assessing information"
- "3.1.1. Information materiality"
source_locator:
source_capture_id: PP-CAP-000001
pdf_page: 7
bounding_boxes: []
character_range:
start: 14820
end: 15430
content_digest: sha256:...
provenance: {}
The AST is a transformed representation and SHALL not replace the source capture.
The AST MAY project reviewed MDX, Source Fragments, chunks, indexes, assertions, citations and comparison views.
16. MDX Projection
Reviewed MDX SHALL be treated as a governed projection of the AST.
The MDX SHALL preserve Source Object, Source Capture, AST version, transformation, reviewer, corrections, unresolved Findings, digest and provenance.
MDX SHALL not silently replace the source artefact.
17. Source Fragment Generation
Source Fragments SHALL follow stable source evidence boundaries.
For regulatory material, candidates SHOULD include numbered paragraphs, lettered subparagraphs, Application Requirements, AR subparagraphs, definitions, tables, table rows, Appendix items, notes and footnotes.
source_fragment:
fragment_id: PP-SF-000023
fragment_version: 1.0.0
document_id: PP-KD-000001
document_version: 1.0.0
fragment_type: numbered_paragraph
source_identifier:
label: "23"
locator:
pdf_page: 7
heading_path:
- "3. Double materiality"
- "3.1.1. Information materiality"
content_digest: sha256:...
rights_record_ids:
- PP-RGT-000001
provenance: {}
Stable numbered fragments SHOULD be generated automatically.
Curator review SHOULD focus on ambiguous boundaries, paragraphs spanning pages, malformed tables, nested ARs, missing identifiers, duplicate extraction and source errors.
18. Content Classification
Content classes MAY include normative, application requirement, guidance, definitions, transitional, appendix, informative, example, note, citation, table, formula, metadata and non-content.
ZARA SHOULD propose the class. Curators SHOULD verify by exception.
Classification SHALL not itself establish legal force.
19. Chunk Planning
Pergamum Pulse SHALL propose Retrieval Chunks.
A numbered paragraph SHOULD normally become a Source Fragment. It MAY or MAY NOT be an optimal Retrieval Chunk.
chunk_plan:
chunk_plan_id: PP-CHPL-000001
chunk_plan_version: 1.0.0
document_id: PP-KD-000001
document_version: 1.0.0
chunking_profile:
profile_id: PP-CHP-REGULATORY-RAG-001
profile_version: 1.0.0
proposed_chunks:
- PP-CHK-000001
- PP-CHK-000002
proposed_by:
actor_type: agent
actor_id: ZARA
agent_profile: pp.chunk-planning.regulatory
status: awaiting_review
provenance: {}
A regulatory Chunking Profile SHOULD preserve numbered paragraphs, AR boundaries, heading context, normative verbs, conditions, exceptions, internal references, defined terms, tables, formulas, units and token budgets.
A chunk MAY contain one substantial paragraph, several short related paragraphs, a paragraph and supporting AR, a heading and related fragments or a coherent table segment.
The same document MAY have separate plans for regulatory evidence, general RAG, legal review, change detection, Academy learning and dataset extraction.
20. Chunk Review Interaction
The UI SHOULD present proposed boundaries rather than empty manual selection.
Recommended visual states are grey for not evaluated, blue for proposed, green for approved, amber for review required, red for rejected or defective and purple for manually overridden.
The seal interaction SHOULD mean: Review and approve this proposed chunk boundary.
The curator SHALL be able to approve, reject, split, merge, move boundary, exclude, reclassify, change chunk purpose, request reprocessing and comment.
Bulk approval SHOULD support low-risk proposals and exception-only review.
21. Chunk Position Semantics
position:
sequence: 10
offset:
type: unicode_code_point
start: 2400
end: 3910
structural_locator:
heading_path:
- "3. Double materiality"
- "3.1.1. Information materiality"
paragraph_range:
start: 22
end: 24
sequence is the ordinal position of the chunk within the output of the identified Chunking Profile for the identified source version.
It SHALL NOT mean importance, authority, evidence strength, paragraph number, page number or priority.
Chunks SHOULD preserve both structural addressing and machine offsets.
22. Semantic Review Stage
After structural admission and chunk approval, ZARA MAY propose Knowledge Assertions, Knowledge Entities, relationships, definitions, requirements, permissions, prohibitions, capabilities, values, thresholds, formulas, applicability and conflicts.
A competent Domain Curator SHALL verify high-impact semantic proposals.
Every assertion candidate SHALL link to Source Fragments.
Draft regulatory knowledge SHALL preserve draft status and SHALL not silently replace active final requirements.
Approved semantic objects SHALL pass the applicable HECATE Profile.
23. Change Intelligence Stage
Pergamum Pulse MAY compare admitted knowledge to prior source versions, current ZAYAZ Registries, Requirements, Controls, Evidence Requirements, schemas, validators, workflows, reports, Publications, Computation Hub models and Origin profiles.
ZARA MAY propose semantic differences, affected objects, affected Modules, affected Components, change severity, draft Change Candidates, tests and documentation updates.
The Regulatory Architect SHALL determine whether to ignore, monitor, create an implementation study, prepare a draft projection, create a Change Candidate, wait for final source or escalate constitutionally.
Draft external knowledge SHOULD default to draft or preview projection only, with production activation prohibited.
24. Final Admission Stage
Final admission SHALL require source identity approval, sufficient Rights Record, accepted AST, complete structural inventory, resolved material extraction conflicts, generated Source Fragments, approved required chunk plan, completed semantic review where required, valid HECATE receipts, resolved or excepted blocking Findings, complete provenance, Module lineage, Component lineage and an Admission Decision.
Admission SHALL not mean public Publication, certification, Registry activation or Runtime activation.
25. Roles and Competence
| Role | Responsibility |
|---|---|
| Source Submitter | Provides source and intended purpose |
| Knowledge Curator | Approves source identity and general metadata |
| Rights Curator | Approves processing and Publication rights |
| Document Curator | Reconciles extraction and canonical AST |
| Retrieval Curator | Approves Chunking Profiles and chunk plans |
| Domain Curator | Reviews subject-matter meaning |
| Regulatory Architect | Reviews impact and Change Candidates |
| Knowledge Governance Approver | Admits governed knowledge |
| HECATE Operator | Maintains validation service without semantic authority |
| Platform Administrator | Maintains workflow and permissions |
Stages MAY require competence records.
High-impact cases SHOULD prohibit submitter as final approver, proposal author as sole approver, uncertain rights proposal self-approval, Agent output entering canonical state directly and HECATE pass replacing governance approval.
Four-eyes approval SHOULD apply to rights expansion, public redistribution, external AI processing, legal interpretation, global promotion, constitutional impact and production Change Candidates.
26. Assignment and Work Management
Workflow tasks SHALL support assignment, reassignment, claiming, delegation, due date, priority, escalation, rework, blocking dependency, parallel review, comments, attachments, mentions and notifications.
workflow_task:
task_id: PP-WFT-000001
case_id: PP-IC-000001
stage: rights_review
task_type: review_rights_record
required_role: rights_curator
assigned_to: actor.example
status: in_progress
priority: high
due_at: 2026-08-10T00:00:00Z
required_fields:
- copyright_holder
- permission_basis
- embedding_permission
- external_processing_permission
gate:
hecate_profile: PP-HCP-RIGHTS-ADMISSION-001
provenance: {}
27. Approval Actions
Every proposed field, object or stage SHALL support accept, modify, reject, escalate, defer, request evidence and request reprocessing.
Acceptance SHALL preserve proposal and evidence.
Modification SHALL preserve proposed value, accepted value, difference, curator, rationale and evidence.
Rejection SHALL preserve reason and effect.
Escalation SHALL identify escalation type, required Role, blocking status, due date and evidence.
28. Database Architecture
Pergamum Pulse SHOULD use a dedicated logical database namespace.
Database: zayaz
Schemas:
pp
hecate
audit
The pp schema SHOULD remain service-extractable if Pergamum Pulse later moves to a separate database.
Use UUIDv7 or ULID for internal relational identity and governed PP-... identifiers for external identity.
Do not use MAX(id) + 1.
Identifier allocation SHALL use PostgreSQL sequences, atomic update-and-return or distributed block allocation.
Core table families SHOULD include:
pp.source_series
pp.source_object
pp.source_object_version
pp.source_capture
pp.source_artifact
pp.transformation
pp.extraction_candidate
pp.extraction_conflict
pp.rights_record
pp.rights_permission
pp.rights_evidence
pp.rights_decision
pp.knowledge_document
pp.knowledge_document_version
pp.document_node
pp.knowledge_section
pp.source_fragment
pp.knowledge_bundle
pp.knowledge_entity
pp.knowledge_assertion
pp.relationship
pp.mapping_object
pp.crosswalk
pp.conflict_set
pp.chunking_profile
pp.chunk_plan
pp.knowledge_chunk
pp.embedding_profile
pp.embedding_record
pp.retrieval_index
pp.ingest_case
pp.ingest_profile
pp.workflow_instance
pp.workflow_stage
pp.workflow_task
pp.assignment
pp.proposal
pp.proposal_field
pp.review_decision
pp.approval
pp.comment
pp.checklist_item
pp.change_candidate
pp.change_decision
pp.projection_manifest
hecate.validation_request
hecate.validation_execution
hecate.finding
hecate.conformance_receipt
pp.event
pp.outbox_event
pp.historical_record
audit.access_event
audit.change_event
Use relational columns for identity, version, state, tenant, white-label operator, valid time, transaction time, foreign keys, assignments and approvals.
Use JSONB for extensible metadata, source-specific payloads, extraction details, Agent explanations and Profile-specific fields.
Use object storage for source artefacts and large projections, graph projection for semantic traversal, vector storage for embeddings and an event log for lifecycle history.
PostgreSQL SHALL remain the system of record for identity and governance state.
29. Proposal Store
Agent outputs SHALL enter a Proposal Store rather than canonical tables directly.
ZARA Output
↓
Proposal Store
↓
HECATE Validation
↓
Curator Decision
↓
Canonical Write Service
proposal:
proposal_id: PP-PROP-000001
case_id: PP-IC-000001
proposal_type: field_value
target_object_type: source_object
target_field: publisher
proposed_value: EFRAG
evidence:
- PP-PEV-000001
confidence:
score: 0.998
method: calibrated
generated_by:
actor_type: agent
actor_id: ZARA
profile: pp.source-metadata-extraction
version: 1.0.0
status: awaiting_review
provenance: {}
Proposal states MAY include generated, validated, awaiting review, accepted, modified, rejected, escalated, expired and superseded.
Only the Canonical Write Service SHALL create or update approved canonical records.
30. Services and Agents
Initial deterministic services SHOULD include Source Capture Service, Identifier Service, File Integrity Service, PDF Extraction Service, Layout Extraction Service, Table Extraction Service, AST Builder, Fragment Generator, Chunk Builder, Embedding Service, Index Builder, Workflow Service, Canonical Write Service, Event Publisher and HECATE Client.
Initial specialised ZARA Agents SHOULD include Source Intake Agent, Source Classification Agent, Document Reconciliation Agent, Rights Triage Agent, Structural Analysis Agent, Chunk Planning Agent, Assertion Extraction Agent, Semantic Mapping Agent, Change Intelligence Agent and Curator Explanation Agent.
Agents SHALL NOT allocate identifiers independently, write canonical records directly, approve rights, approve legal interpretation, approve equivalence, approve Change Candidates, admit knowledge, publish or activate.
31. Orchestration
The pipeline SHOULD be event-driven and workflow-orchestrated.
Workflow Orchestrator
├── deterministic worker: capture
├── deterministic worker: extract
├── deterministic worker: build AST
├── ZARA Agent: classify source
├── ZARA Agent: reconcile extraction
├── HECATE: validate structure
├── Human Task: review exceptions
├── ZARA Agent: propose chunks
├── Human Task: approve chunk plan
└── Canonical Write Service: admit objects
Every worker SHALL support idempotency keys.
Retries SHALL preserve original execution, retry count, failure, output and provenance.
Failed jobs SHALL enter a governed dead-letter queue.
32. Event Model
Events SHOULD use an outbox pattern.
Core events MAY include:
pp.ingest_case.created
pp.source_capture.completed
pp.source_identity.proposed
pp.source_identity.approved
pp.rights_review.requested
pp.rights_record.approved
pp.transformation.completed
pp.extraction_conflict.detected
pp.document_ast.proposed
pp.document_ast.approved
pp.source_fragments.generated
pp.chunk_plan.proposed
pp.chunk_plan.approved
pp.assertions.proposed
pp.assertions.approved
pp.change_candidate.proposed
pp.hecate.receipt.issued
pp.knowledge_document.admitted
pp.case.blocked
pp.case.completed
Every event SHALL identify event identity, event type, aggregate identity, aggregate version, case, actor, tenant, white-label operator, valid time, transaction time, payload digest and provenance.
33. Permissions and Security
Permissions SHALL be Role-, stage-, object- and scope-aware.
Permission examples include view source, view restricted source, edit proposal, approve source identity, approve rights, approve AST, approve chunk plan, approve assertion, approve admission, create exception, export source, publish and promote tenant knowledge.
Tenant and white-label boundaries SHALL apply to source artefacts, cases, proposals, comments, tasks, ASTs, fragments, chunks, embeddings, Findings, receipts, exports and logs.
Restricted rights evidence, legal advice and confidential source material SHOULD support field- or object-level protection.
34. HECATE Stage Gates
Every workflow stage SHOULD have a HECATE Profile.
Examples include Intake Gate, Source Identity Gate, Rights Gate, Extraction Gate, Structural Gate, Chunk Plan Gate, Semantic Admission Gate and Final Knowledge Admission Gate.
Gate outcomes MAY include pass, pass with warnings, conditional pass, fail, blocked and unable to determine.
A HECATE pass SHALL not replace required human approval.
Exceptions SHALL be explicit, authorised, scoped, time-bound, visible and historically preserved.
35. Rework and Exception Handling
A stage MAY be returned for rework.
The system SHALL preserve originating stage, destination stage, reason, Findings, actor, due date, affected approvals and provenance.
A material upstream change SHALL identify which downstream approvals and receipts expire.
Dedicated exception queues SHOULD exist for rights uncertainty, source identity conflict, extraction conflict, structural anomaly, chunk anomaly, semantic conflict, HECATE failure and tenant scope conflict.
36. Notifications
The system SHOULD notify actors about new assignment, approaching due date, overdue task, blocking Finding, escalation, requested review, rework, approval, rejection, rights expiry, source supersession and receipt expiry.
Notifications SHALL not expose restricted source content.
37. Audit and Historical Reconstruction
The system SHALL reconstruct source submission, capture, transformations, proposals, evidence, curator actions, HECATE executions, exceptions, approvals, canonical writes, admitted knowledge, corrections, supersession and withdrawal.
Every proposal and Decision SHALL remain historically available.
The system SHALL preserve source, AST, MDX, fragment, chunk-plan, chunk and assertion versions.
The system SHOULD support exact workflow replay, exact extraction replay where tools are preserved, semantic replay under later Profiles and receipt verification.
38. User Experience Principles
The Workbench SHALL show evidence next to proposals, minimise unnecessary clicks, support grouped low-risk acceptance, surface uncertainty, preserve conflicts, show blocking reasons, separate proposal from approval, show current stage and next action, preserve source context, support keyboard operation, support large documents and avoid forcing curators through irrelevant content.
The default view SHOULD show proposed value, confidence, evidence, status and required action.
Advanced details MAY show model, prompt, extraction method, digests, provenance graph and Rule results.
ZARA SHOULD explain why a proposal was made.
The interface SHALL not visually present an unapproved proposal as canonical.
39. Regulatory Draft Pilot
The first regulatory pilot SHOULD use an EFRAG draft ESRS document.
pilot:
source_class: regulatory_standard_draft
publisher: EFRAG
source_format: PDF
transformation_candidates:
- native_pdf_text
- layout_extraction
- table_extraction
- third_party_markdown
production_activation: prohibited
Milestone 1 SHALL deliver Source Series, Source Object, Source Capture, preliminary Rights Record and file digests.
Milestone 2 SHALL deliver extraction candidates, extraction conflicts, reviewed AST, reviewed MDX, structural inventory and HECATE structural receipt.
Milestone 3 SHALL deliver Source Fragments, regulatory Chunking Profile, chunk plan, approved chunks, lexical index and embedding eligibility Decision.
Milestone 4 SHALL deliver reviewed assertion sample, evidence relationships, entities, conflicts and HECATE assertion receipts.
Milestone 5 SHALL deliver comparison baseline, semantic differences, affected Modules, affected Components, draft Change Candidates and no production activation.
40. Viroway-Authored Pilot Profile
The next pilot SHOULD cover a Viroway-authored ZAYAZ document.
source_class: viroway_internal_document
copyright_holder: Viroway Ltd
ownership: full
internal_processing: permitted
embedding: permitted_internal
external_processing: governed_by_provider_policy
publication: separate_approval_required
internal_authority: governed_by_document_status
The Viroway Profile MAY automatically accept copyright holder, standard attribution, internal archiving, internal parsing, internal indexing and internal embeddings where the document originated from an approved Viroway repository and no contrary metadata exists.
Ownership SHALL not automatically create constitutional authority.
41. API Boundaries
Initial APIs SHOULD include:
POST /pp/ingest-cases
GET /pp/ingest-cases/{id}
POST /pp/ingest-cases/{id}/artifacts
POST /pp/ingest-cases/{id}/proposals
POST /pp/proposals/{id}/accept
POST /pp/proposals/{id}/modify
POST /pp/proposals/{id}/reject
POST /pp/proposals/{id}/escalate
POST /pp/ingest-cases/{id}/advance
POST /pp/ingest-cases/{id}/return
GET /pp/ingest-cases/{id}/workbench
GET /pp/ingest-cases/{id}/findings
POST /pp/chunk-plans/{id}/approve
POST /pp/chunk-plans/{id}/override
POST /pp/hecate/validate
POST /pp/knowledge-documents/{id}/admit
Canonical writes SHALL require authenticated actor, permission, valid stage, required Gate, expected object version, idempotency key and provenance.
42. Non-Functional Requirements
The Workbench SHOULD support documents exceeding 1,000 pages, thousands of Source Fragments, multiple extraction candidates, concurrent reviewers, resumable processing, partial stage loading, optimistic concurrency, full audit, tenant isolation, white-label theming, accessible keyboard navigation, configurable retention and EU data residency where required.
Concurrent edits SHALL use version checks, locks for high-risk objects, conflict resolution and no silent last-write-wins.
43. Implementation Sequence
Recommended implementation order:
ppdatabase schema and identifiers;- Ingest Case and Workflow Service;
- immutable Source Capture Service;
- source identity proposals;
- Kanban board;
- Case Workbench shell;
- PDF viewer and source locator;
- extraction candidate storage;
- AST Builder;
- Source Fragment Generator;
- Proposal Store;
- HECATE stage Gates;
- Chunk Planning Agent;
- chunk-review interface;
- semantic proposal workflow;
- change-intelligence workflow;
- final admission;
- historical replay.
44. Architectural Decisions
The following decisions are established:
- The source artefact is immutable.
- Converted Markdown is a Transformation candidate.
- The canonical machine representation is a document AST.
- MDX is a governed projection.
- Source Fragments follow source evidence boundaries.
- Retrieval Chunks are purpose-specific derived objects.
- Pergamum Pulse proposes chunk plans.
- Curators review by exception.
- ZARA writes proposals, not canonical state.
- HECATE validates objective conformance only.
- Kanban manages work; the Case Workbench handles content.
- Workflow is Profile-driven.
- Different Roles may approve different stages.
- PostgreSQL is the governance system of record.
- Graph and vector stores are governed projections.
- Internal UUIDs and governed public identifiers are separate.
- Canonical writes pass through a dedicated service.
- Every stage preserves Module and Component lineage.
- Admission is separate from Publication and activation.
- Historical reconstruction is mandatory.
45. Conformance Requirements
An implementation conforms where it:
- creates an Ingest Case for every governed intake;
- captures source artefacts immutably;
- computes and stores source digests;
- separates deterministic processing from semantic proposal;
- separates AI proposal from canonical write;
- represents proposals as first-class objects;
- preserves proposal evidence, confidence and method;
- supports accept, modify, reject and escalate;
- preserves curator rationale and provenance;
- supports grouped review by exception;
- uses configuration-driven Ingest Profiles;
- supports stage-specific required fields;
- supports stage-specific Roles;
- supports HECATE stage Gates;
- prevents Kanban drag-and-drop from bypassing Gates;
- separates Kanban workload management from detailed workbench review;
- provides source, normalised and review panes;
- displays evidence beside proposals;
- preserves extraction candidates independently;
- treats converted Markdown as a Transformation;
- preserves the source artefact as source of record;
- supports a canonical document AST;
- preserves AST node identity, type, source locator and digest;
- projects reviewed MDX from governed structure;
- generates stable Source Fragments;
- preserves fragment rights and provenance;
- distinguishes Source Fragments from Retrieval Chunks;
- allows multiple Chunking Profiles per document;
- has Pergamum Pulse propose chunk boundaries;
- supports curator split, merge, move, exclude and override;
- preserves chunk sequence and explicit offset type;
- uses dual structural and machine addressing;
- supports content-class proposals;
- prevents content class from automatically creating legal force;
- supports specialised ZARA Agents;
- prevents Agents from approving or writing canonical records directly;
- uses deterministic services for mechanical processing;
- makes workers idempotent and retry-safe;
- records worker and Agent versions;
- uses an atomic identifier service;
- does not allocate identifiers through
MAX(id) + 1; - separates internal database identity from public identity;
- stores governance state relationally;
- preserves extensible metadata without collapsing all data into JSONB;
- uses object storage for large artefacts;
- treats graph and vector stores as projections;
- uses PostgreSQL as identity and governance system of record;
- supports assignment, delegation, due dates and escalation;
- supports separation of duties;
- supports four-eyes approval for high-impact Decisions;
- applies tenant and white-label isolation throughout;
- protects restricted rights and legal evidence;
- supports rework and approval invalidation;
- preserves all historical proposals and Decisions;
- emits governed lifecycle events;
- uses an outbox or equivalent reliable event pattern;
- supports optimistic concurrency and conflict detection;
- preserves Module lineage;
- preserves Component lineage;
- supports HECATE receipts and Findings;
- prevents HECATE passes from replacing authority;
- preserves exceptions and waivers;
- separates knowledge admission from Publication;
- separates knowledge admission from projection;
- separates knowledge admission from activation;
- supports regulatory draft protection;
- supports Viroway-authored source Profiles;
- supports exact historical reconstruction;
- supports future semantic replay without rewriting history;
- exposes clear visual distinction among proposed, approved, rejected and overridden states.
46. Foundational Principle
Pergamum Pulse SHALL not make curators perform work that deterministic services and governed AI can perform more consistently, quickly and transparently.
It SHALL not allow automation to absorb human, legal, semantic, constitutional or operational authority.
The source remains immutable. Extraction remains traceable. The AST remains a representation. MDX remains a projection. Fragments remain evidence boundaries. Chunks remain purpose-specific retrieval objects. ZARA remains a proposal engine. HECATE remains a validator. Curators and authorised governance bodies remain accountable for Decisions.
The Ingest Workbench therefore exists to make every stage visible: what the system observed, what it extracted, what it inferred, what it proposed, what evidence supports it, what the curator accepted, what HECATE validated and what was ultimately admitted.
By combining deterministic services, specialised Agents, evidence-first review, configurable workflow, separation of duties, tenant isolation and historical replay, Pergamum Pulse can scale from a single document pilot to a global knowledge-governance infrastructure without sacrificing precision, accountability or trust.