9 Commits
Author SHA1 Message Date
zemion 7a7654cc0f feat(datasources): govern approvals and retention
Module Package Release / publish-packages (push) Successful in 11s
2026-08-22 19:37:44 +02:00
zemion b54d1919e4 feat(datasources): enforce policy-aware data visibility
Module Package Release / publish-packages (push) Successful in 11s
2026-08-21 20:31:52 +02:00
zemion 97ca670bfe feat(datasources): publish durable artifact outputs
Module Package Release / publish-packages (push) Successful in 12s
2026-08-21 17:36:19 +02:00
zemion f03497bdaf feat(datasources): add governed DSAR coverage 2026-08-21 03:19:42 +02:00
zemion 1bb03c0f91 refactor(webui): adopt semantic workspace actions 2026-08-19 18:47:45 +02:00
zemion 45cc3cdc7a feat: classify datasources product area 2026-08-18 21:32:50 +02:00
zemion 32a32c9906 Adopt shared WebUI structural primitives 2026-08-18 13:17:31 +02:00
zemion c16bb51aa3 Adopt shared WebUI layout primitives 2026-08-18 10:42:51 +02:00
zemion 492306f6a7 Add permission-aware datasource search source 2026-08-06 16:06:17 +02:00
34 changed files with 7045 additions and 500 deletions
+33 -4
View File
@@ -22,8 +22,11 @@ explicit frozen states, previews, retirement, and atomic producer publication.
Origin discovery retains each provider's source mode, structured health, and Origin discovery retains each provider's source mode, structured health, and
declared pushdown support. Live previews preserve the provider's effective row, declared pushdown support. Live previews preserve the provider's effective row,
serialized-byte, and elapsed-time limits and its redacted diagnostics. serialized-byte, and elapsed-time limits and its redacted diagnostics.
Producer modules can append a bounded tabular result or create a new static Producer modules can append a bounded inline tabular result or pin a larger
datasource through an idempotent capability. The publication ledger retains the durable artifact through the same idempotent capability. Artifact references
declare a backend, locator, SHA-256 checksum, schema, fingerprint, row and byte
counts; the installed provider verifies integrity and serves bounded reads.
The publication ledger retains the
producer run, output materialization, provenance, and replay identity. On producer run, output materialization, provenance, and replay identity. On
PostgreSQL, a transaction-scoped advisory lock serializes each tenant, producer, PostgreSQL, a transaction-scoped advisory lock serializes each tenant, producer,
and idempotency identity before any output side effect, so retries from multiple and idempotency identity before any output side effect, so retries from multiple
@@ -35,9 +38,35 @@ promotion preserves the exact policy hash and validation result in immutable
materialization provenance. The supported contract is documented in materialization provenance. The supported contract is documented in
[docs/QUALITY_POLICY.md](docs/QUALITY_POLICY.md). [docs/QUALITY_POLICY.md](docs/QUALITY_POLICY.md).
The same lifecycle contract can require an attributable, separated approval
quorum before a stage or cached refresh becomes current. Versioned retention
rules cover transient stages, ordinary materializations, and frozen evidence.
Retention remains an explicit preview-and-apply operation: current revisions,
legal holds, pending approvals, and publication evidence are blocked, while
every decision, promotion, and disposition is preserved in hash-chained local
evidence even when the optional Audit, Access, or Policy modules are absent.
The contracts already model database, HTTP/REST, directory, file, feed, The contracts already model database, HTTP/REST, directory, file, feed,
document, binary, directory, and stream sources so providers can be added document, binary, directory, and stream sources so providers can be added
without changing consumers. Larger durable artifact-backed publications remain without changing consumers. Storage modules contribute artifact backends
a later storage-provider slice. through the provider-neutral `datasources.artifactBackends` capability; the
Datasources module never imports their internals.
See [docs/CONCEPT.md](docs/CONCEPT.md) for ownership and lifecycle details. See [docs/CONCEPT.md](docs/CONCEPT.md) for ownership and lifecycle details.
## Data-subject requests
Datasources publishes `privacy.dsar.datasources` for exact catalogue,
governance-reference, materialization, payload, stage, publication, and
lifecycle-evidence
references and for minimized operator attribution. It never exports connector
references, locators, credentials, arbitrary rows, schemas, validation
samples, metadata, provenance bodies, checkpoints, replay material, or hashes.
The module does not guess subject identity by scanning schema-dependent tabular
payloads; the authoritative source module locates and corrects those facts.
Unpromoted stages and unreferenced payloads can be deleted idempotently.
Published or referenced state, immutable materializations, lifecycle evidence,
governance evidence,
holds, and operator attribution require data-steward review. Dataflow and
Reporting derivatives must be refreshed after the source correction.
+22
View File
@@ -88,6 +88,21 @@ The first slice stores bounded tabular JSON/CSV stages. Future providers may
stage file references, object-store blobs, directory snapshots, or streaming stage file references, object-store blobs, directory snapshots, or streaming
checkpoints through the same lifecycle contract. checkpoints through the same lifecycle contract.
Valid stages may be subject to a versioned local approval policy. Required
quorums keep the stage non-consumable, enforce distinct attributable actors and
optional creator/approver separation, and bind every decision to exact policy
and subject hashes. Cached refreshes use this same staging path when approval is
required. An installation without separate Access or Policy modules retains
safe local scope and policy defaults; optional modules do not manufacture an
approval from configuration metadata.
Retention policy separately defines durations for stages, ordinary
materializations, and frozen evidence. Administrators receive a deterministic
plan before applying any disposition. Holds, current state, pending decisions,
and publication evidence are blockers. Payload disposal preserves a minimized
materialization tombstone and hash-chained lifecycle evidence, so chronological
authority and policy proof survive the configured content-retention boundary.
## Consumer Contract ## Consumer Contract
Consumers request: Consumers request:
@@ -104,6 +119,13 @@ Consumers should be able to request the governance explanation and dependency
impact separately from row access. Seeing catalogue metadata must not imply impact separately from row access. Seeing catalogue metadata must not imply
permission to read protected data. permission to read protected data.
When Search is installed, Datasources contributes catalogue entries as a native
search source. Only the stable catalogue identity, display name, description,
mode, shape, lifecycle state, classification, publication state, and authority
mode are indexed. Every result is re-authorized against the current tenant and
catalogue-read permission. Rows, schemas, connector references, credentials,
arbitrary metadata, and provenance remain outside the derived search index.
## Next Providers ## Next Providers
Connector providers should cover: Connector providers should cover:
+76 -3
View File
@@ -96,6 +96,79 @@ This makes concurrent retries from separate API or worker nodes converge on the
same publication and materialization rather than relying on a late uniqueness same publication and materialization rather than relying on a late uniqueness
failure after output rows have already been persisted. failure after output rows have already been persisted.
Approval authority, approval expiry, and retention/deletion execution remain ## Durable artifact publications
separate work under `govoplan-datasources#2`. Until those contracts are added,
no JSON flag is treated as an approval and no stage is deleted automatically. Outputs larger than the inline row and byte limits use an immutable artifact
reference. The reference pins its backend and locator together with SHA-256
checksum, schema, datasource fingerprint, row count, byte count, media type,
and optional resume checkpoint. Datasources persists that reference as the
materialization payload and asks the installed Core-contract artifact backend
to verify it before creating catalogue state. Reads remain bounded and are
re-authorized by Datasources before reaching the backend.
Schema rules are evaluated by Datasources. Content-level rules such as
uniqueness or range require producer evidence bound to the exact payload
checksum and current quality-policy hash, including all evaluated rule IDs.
Missing or explicitly deferred evidence produces a `review_required`
publication and an immutable, addressable materialization, but it does not
replace the Datasource's current state. Valid warnings produce
`published_with_warnings`; failed evidence blocks the publication without a
catalogue side effect. These terminal states are preserved for Dataflow and
Workflow handoffs instead of being collapsed into generic success.
## Promotion approval
Approval is a separate deterministic contract in
`governance.approval_policy`:
```json
{
"version": "monthly-promotion-v2",
"required": true,
"required_approvals": 2,
"separation_of_duties": true,
"expires_after_hours": 72,
"policy_ref": "policy:monthly-register-promotion"
}
```
A valid stage enters `awaiting_approval` instead of `ready`. Each decision is
bound to the stage fingerprint, quality-policy hash, normalized approval-policy
hash, actor authority, reason, and expiry. Actors are distinct and the stage
creator cannot approve when separation of duties is enabled. The exact quorum
snapshot and hash-chained evidence are copied into materialization provenance.
A changed target policy or staged subject requires a new stage; request JSON
cannot claim that an approval occurred.
Cached origins use `POST /datasources/{id}/refresh/stage` when approval is
required. Direct refresh then fails closed, so connector content cannot become
current before the staged validation and approval quorum succeed.
## Retention
Retention is independently configured in `governance.retention_policy`:
```json
{
"version": "register-retention-v3",
"enabled": true,
"stage_days": 30,
"materialization_days": 365,
"frozen_evidence_days": 3650,
"policy_ref": "policy:register-retention"
}
```
Durations are optional; omitting a class retains it indefinitely. Empty or
disabled local policy never deletes content. An administrator first requests a
read-only plan. The plan includes all due targets, policy versions and hashes,
eligibility dates, blockers, and one hash over the complete preview. Applying
retention accepts only targets from the unchanged plan.
Pending approvals, current materializations, legal holds, and materializations
referenced by producer publications are blocked. Eligible stages are deleted.
Eligible materialization payload rows are purged while the revision retains its
schema, provenance, row/byte counts, payload checksum, disposition metadata,
and immutable hash-chained lifecycle evidence. Frozen evidence is eligible only
when `frozen_evidence_days` is explicitly configured. No hidden scheduler or
arbitrary Policy/Access flag performs deletion.
+2 -2
View File
@@ -4,13 +4,13 @@ build-backend = "setuptools.build_meta"
[project] [project]
name = "govoplan-datasources" name = "govoplan-datasources"
version = "0.1.18" version = "0.1.21"
description = "Governed datasource catalogue, staging, and materialization lifecycle for GovOPlaN." description = "Governed datasource catalogue, staging, and materialization lifecycle for GovOPlaN."
readme = "README.md" readme = "README.md"
requires-python = ">=3.12" requires-python = ">=3.12"
license = "AGPL-3.0-or-later" license = "AGPL-3.0-or-later"
authors = [{ name = "GovOPlaN" }] authors = [{ name = "GovOPlaN" }]
dependencies = ["govoplan-core>=0.1.18"] dependencies = ["govoplan-core>=0.1.20"]
[tool.setuptools.packages.find] [tool.setuptools.packages.find]
where = ["src"] where = ["src"]
+1 -1
View File
@@ -1,3 +1,3 @@
"""GovOPlaN Datasources module.""" """GovOPlaN Datasources module."""
__version__ = "0.1.18" __version__ = "0.1.21"
@@ -109,6 +109,10 @@ class DatasourceRecord(Base, TimestampMixin):
) )
privacy_profile_ref: Mapped[str | None] = mapped_column(String(500), nullable=True) privacy_profile_ref: Mapped[str | None] = mapped_column(String(500), nullable=True)
retention_policy_ref: Mapped[str | None] = mapped_column(String(500), nullable=True) retention_policy_ref: Mapped[str | None] = mapped_column(String(500), nullable=True)
access_policy_ref: Mapped[str | None] = mapped_column(String(500), nullable=True)
visibility_policy: Mapped[dict[str, Any]] = mapped_column(
JSON, default=dict, nullable=False
)
hold_refs: Mapped[list[str]] = mapped_column(JSON, default=list, nullable=False) hold_refs: Mapped[list[str]] = mapped_column(JSON, default=list, nullable=False)
publication_state: Mapped[str] = mapped_column( publication_state: Mapped[str] = mapped_column(
String(50), default="draft", nullable=False, index=True String(50), default="draft", nullable=False, index=True
@@ -122,6 +126,12 @@ class DatasourceRecord(Base, TimestampMixin):
quality_policy: Mapped[dict[str, Any]] = mapped_column( quality_policy: Mapped[dict[str, Any]] = mapped_column(
JSON, default=dict, nullable=False JSON, default=dict, nullable=False
) )
approval_policy: Mapped[dict[str, Any]] = mapped_column(
JSON, default=dict, nullable=False
)
retention_policy: Mapped[dict[str, Any]] = mapped_column(
JSON, default=dict, nullable=False
)
known_limits: Mapped[list[str]] = mapped_column(JSON, default=list, nullable=False) known_limits: Mapped[list[str]] = mapped_column(JSON, default=list, nullable=False)
correction_procedure_ref: Mapped[str | None] = mapped_column( correction_procedure_ref: Mapped[str | None] = mapped_column(
String(500), nullable=True String(500), nullable=True
@@ -262,6 +272,15 @@ class DatasourceMaterializationRecord(Base, TimestampMixin):
default=dict, default=dict,
nullable=False, nullable=False,
) )
disposed_at: Mapped[datetime | None] = mapped_column(
DateTime(timezone=True), nullable=True, index=True
)
disposition_: Mapped[dict[str, Any]] = mapped_column(
"disposition",
JSON,
default=dict,
nullable=False,
)
created_by: Mapped[str | None] = mapped_column(String(255), nullable=True, index=True) created_by: Mapped[str | None] = mapped_column(String(255), nullable=True, index=True)
datasource: Mapped[DatasourceRecord] = relationship(back_populates="materializations") datasource: Mapped[DatasourceRecord] = relationship(back_populates="materializations")
@@ -413,6 +432,12 @@ class DatasourceStageRecord(Base, TimestampMixin):
default=dict, default=dict,
nullable=False, nullable=False,
) )
approval_: Mapped[dict[str, Any]] = mapped_column(
"approval",
JSON,
default=dict,
nullable=False,
)
promoted_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True), nullable=True) promoted_at: Mapped[datetime | None] = mapped_column(DateTime(timezone=True), nullable=True)
promoted_materialization_id: Mapped[str | None] = mapped_column( promoted_materialization_id: Mapped[str | None] = mapped_column(
String(36), String(36),
@@ -479,7 +504,49 @@ class DatasourcePublicationRecord(Base, TimestampMixin):
) )
class DatasourceLifecycleEvidenceRecord(Base):
__tablename__ = "datasource_lifecycle_evidence"
__table_args__ = (
UniqueConstraint(
"tenant_id",
"event_hash",
name="uq_datasource_lifecycle_evidence_hash",
),
Index(
"ix_datasource_lifecycle_evidence_subject",
"tenant_id",
"subject_ref",
"occurred_at",
),
Index(
"ix_datasource_lifecycle_evidence_event",
"tenant_id",
"event_type",
"occurred_at",
),
)
id: Mapped[str] = mapped_column(String(36), primary_key=True, default=new_uuid)
tenant_id: Mapped[str] = mapped_column(String(36), nullable=False, index=True)
subject_ref: Mapped[str] = mapped_column(String(160), nullable=False, index=True)
event_type: Mapped[str] = mapped_column(String(80), nullable=False, index=True)
occurred_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), nullable=False)
actor_ref: Mapped[str | None] = mapped_column(String(255), nullable=True, index=True)
policy_version: Mapped[str | None] = mapped_column(String(120), nullable=True)
policy_hash: Mapped[str | None] = mapped_column(String(64), nullable=True)
subject_digest: Mapped[str] = mapped_column(String(64), nullable=False)
previous_event_hash: Mapped[str | None] = mapped_column(String(64), nullable=True)
event_hash: Mapped[str] = mapped_column(String(64), nullable=False, index=True)
details_: Mapped[dict[str, Any]] = mapped_column(
"details",
JSON,
default=dict,
nullable=False,
)
__all__ = [ __all__ = [
"DatasourceLifecycleEvidenceRecord",
"DatasourceMaterializationRecord", "DatasourceMaterializationRecord",
"DatasourcePayloadRecord", "DatasourcePayloadRecord",
"DatasourcePayloadRowRecord", "DatasourcePayloadRowRecord",
@@ -0,0 +1,794 @@
from __future__ import annotations
from collections.abc import Sequence
from dataclasses import dataclass
from datetime import datetime, timezone
from typing import Any
from sqlalchemy import or_
from sqlalchemy.orm import Session
from govoplan_core.core.dsar import (
DsarErasureActionRef,
DsarExecutionResultRef,
DsarRecordRef,
DsarSubjectRef,
dsar_capability_name,
)
from govoplan_datasources.backend.db.models import (
DatasourceGovernanceReferenceRecord,
DatasourceLifecycleEvidenceRecord,
DatasourceMaterializationRecord,
DatasourcePayloadRecord,
DatasourcePublicationRecord,
DatasourceRecord,
DatasourceStageRecord,
)
DATASOURCES_DSAR_CAPABILITY = dsar_capability_name("datasources")
_MAX_RECORDS = 5_000
_CONFLICT = object()
_DIRECT_ALIASES = {
"datasource_id": ("datasources.datasource", "datasources.catalogue"),
"governance_reference_id": ("datasources.governance_reference",),
"materialization_id": ("datasources.materialization",),
"payload_id": ("datasources.payload",),
"stage_id": ("datasources.stage",),
"publication_id": ("datasources.publication",),
"lifecycle_evidence_id": ("datasources.lifecycle_evidence",),
}
_RESOURCE_MODELS = {
"datasource": DatasourceRecord,
"datasource_governance_reference": DatasourceGovernanceReferenceRecord,
"datasource_materialization": DatasourceMaterializationRecord,
"datasource_payload": DatasourcePayloadRecord,
"datasource_stage": DatasourceStageRecord,
"datasource_publication": DatasourcePublicationRecord,
"datasource_lifecycle_evidence": DatasourceLifecycleEvidenceRecord,
}
@dataclass(frozen=True, slots=True)
class _Selectors:
account_id: str | None
identity_id: str | None
membership_id: str | None
direct: dict[str, str]
@property
def actor_ids(self) -> tuple[str, ...]:
return tuple(
value
for value in (self.account_id, self.identity_id, self.membership_id)
if value
)
@dataclass(frozen=True, slots=True)
class _Match:
resource_type: str
row: Any
category: str
class DatasourcesDsarProvider:
provider_id = "datasources"
module_id = "datasources"
def search_subject(
self,
session: object,
*,
tenant_id: str,
subject: DsarSubjectRef,
) -> Sequence[DsarRecordRef]:
db = _session(session)
selectors = _selectors(subject)
if selectors is None or not (selectors.actor_ids or selectors.direct):
return ()
direct = _direct_matches(db, tenant_id=tenant_id, selectors=selectors)
if direct is None:
return ()
if direct:
if selectors.actor_ids and not any(
_correlates(match, selectors.actor_ids) for match in direct
):
return ()
matches = direct
else:
matches = _canonical_matches(
db,
tenant_id=tenant_id,
actor_ids=selectors.actor_ids,
)
records: list[DsarRecordRef] = []
seen: set[tuple[str, str]] = set()
for match in matches:
key = (match.resource_type, str(match.row.id))
if key in seen:
continue
if len(records) >= _MAX_RECORDS:
raise ValueError(
"Datasources DSAR result limit exceeded; narrow the selectors."
)
seen.add(key)
records.append(_record(match))
return tuple(records)
def plan_erasure(
self,
session: object,
*,
tenant_id: str,
subject: DsarSubjectRef,
records: Sequence[DsarRecordRef],
) -> Sequence[DsarErasureActionRef]:
del tenant_id
_session(session)
if _selectors(subject) is None:
raise ValueError("Datasources DSAR subject selectors conflict.")
actions: list[DsarErasureActionRef] = []
for record in records:
_validate_record(record)
if record.category in {
"unpromoted_datasource_stage",
"unreferenced_datasource_payload",
}:
kind = "delete"
executable = True
rationale = (
"Remove transient Datasources content that has not become "
"immutable or referenced lifecycle evidence."
)
elif record.category == "datasource_operator_attribution":
kind = "retain"
executable = False
rationale = record.retention_reason or (
"Institutional data operations remain attributable."
)
else:
kind = "manual_review"
executable = False
rationale = (
"A data steward must correct the authoritative source and "
"review immutable revisions, holds, consumers, and evidence."
)
actions.append(
DsarErasureActionRef(
action_id=(
f"datasources:{kind}:{record.resource_type}:"
f"{record.resource_id}"
),
provider_id=self.provider_id,
module_id=self.module_id,
kind=kind,
resource_type=record.resource_type,
resource_id=record.resource_id,
title=(
f"Delete {record.title}"
if executable
else f"Review {record.title}"
),
rationale=rationale,
executable=executable,
irreversible=executable,
metadata={"record_category": record.category},
)
)
return tuple(actions)
def execute_erasure(
self,
session: object,
*,
tenant_id: str,
subject: DsarSubjectRef,
actions: Sequence[DsarErasureActionRef],
request_id: str,
) -> Sequence[DsarExecutionResultRef]:
db = _session(session)
selectors = _selectors(subject)
if selectors is None:
raise ValueError("Datasources DSAR subject selectors conflict.")
results: list[DsarExecutionResultRef] = []
for action in actions:
_validate_action(action)
if not action.executable:
results.append(
DsarExecutionResultRef(
action_id=action.action_id,
status="blocked",
summary=(
"Use governed source correction and Datasources "
"retention review before changing published state."
),
evidence={"request_id": request_id},
)
)
continue
if action.kind != "delete" or action.resource_type not in {
"datasource_stage",
"datasource_payload",
}:
raise ValueError("Datasources DSAR executable action is unsupported.")
model = _RESOURCE_MODELS[action.resource_type]
row = (
db.query(model)
.filter(model.tenant_id == tenant_id, model.id == action.resource_id)
.with_for_update()
.one_or_none()
)
if row is None:
status = "unchanged"
summary = "Transient Datasources row was already absent."
else:
match = _Match(action.resource_type, row, "execution")
if not (
_directly_targets(selectors, match)
or _correlates(match, selectors.actor_ids)
):
raise ValueError(
"Datasources DSAR action is not corroborated by the subject."
)
_assert_deletable(db, resource_type=action.resource_type, row=row)
db.delete(row)
db.flush()
status = "executed"
summary = "Transient, unreferenced Datasources row removed."
results.append(
DsarExecutionResultRef(
action_id=action.action_id,
status=status,
summary=summary,
evidence={"request_id": request_id},
)
)
return tuple(results)
def _direct_matches(
session: Session,
*,
tenant_id: str,
selectors: _Selectors,
) -> list[_Match] | None:
matches: list[_Match] = []
for selector, value in selectors.direct.items():
if selector == "datasource_id":
datasource = _one(
session,
DatasourceRecord,
tenant_id=tenant_id,
field="id",
value=value.removeprefix("datasource:"),
)
if datasource is None:
return None
current = _datasource_package(
session,
tenant_id=tenant_id,
datasource=datasource,
)
else:
model, resource_type = {
"governance_reference_id": (
DatasourceGovernanceReferenceRecord,
"datasource_governance_reference",
),
"materialization_id": (
DatasourceMaterializationRecord,
"datasource_materialization",
),
"payload_id": (DatasourcePayloadRecord, "datasource_payload"),
"stage_id": (DatasourceStageRecord, "datasource_stage"),
"publication_id": (
DatasourcePublicationRecord,
"datasource_publication",
),
"lifecycle_evidence_id": (
DatasourceLifecycleEvidenceRecord,
"datasource_lifecycle_evidence",
),
}[selector]
row = _one(
session,
model,
tenant_id=tenant_id,
field="id",
value=_strip_prefix(value),
)
if row is None:
return None
current = [
_Match(
resource_type, row, _direct_category(session, resource_type, row)
)
]
matches.extend(current)
if len(matches) > _MAX_RECORDS:
raise ValueError(
"Datasources DSAR result limit exceeded; narrow the selectors."
)
roots = {
root
for match in matches
for root in _root_datasource_ids(session, match)
if root
}
if len(roots) > 1:
return None
return matches
def _datasource_package(
session: Session,
*,
tenant_id: str,
datasource: DatasourceRecord,
) -> list[_Match]:
matches = [_Match("datasource", datasource, "datasource_configuration")]
specs = (
(
DatasourceGovernanceReferenceRecord,
"datasource_governance_reference",
),
(DatasourceMaterializationRecord, "datasource_materialization"),
(DatasourceStageRecord, "datasource_stage"),
(DatasourcePublicationRecord, "datasource_publication"),
)
materializations: list[DatasourceMaterializationRecord] = []
for model, resource_type in specs:
rows = (
session.query(model)
.filter(
model.tenant_id == tenant_id,
(
model.target_datasource_id == datasource.id
if model is DatasourceStageRecord
else model.datasource_id == datasource.id
),
)
.order_by(model.id)
.limit(_MAX_RECORDS + 1)
.all()
)
if model is DatasourceMaterializationRecord:
materializations = rows
matches.extend(
_Match(
resource_type,
row,
(
"datasource_related_stage"
if resource_type == "datasource_stage"
else _direct_category(session, resource_type, row)
),
)
for row in rows
)
payload_ids = {row.payload_id for row in materializations if row.payload_id}
if payload_ids:
payloads = (
session.query(DatasourcePayloadRecord)
.filter(
DatasourcePayloadRecord.tenant_id == tenant_id,
DatasourcePayloadRecord.id.in_(payload_ids),
)
.order_by(DatasourcePayloadRecord.id)
.limit(_MAX_RECORDS + 1)
.all()
)
matches.extend(
_Match(
"datasource_payload",
row,
_direct_category(session, "datasource_payload", row),
)
for row in payloads
)
return matches
def _canonical_matches(
session: Session,
*,
tenant_id: str,
actor_ids: tuple[str, ...],
) -> list[_Match]:
if not actor_ids:
return []
specs = (
(
DatasourceRecord,
or_(
DatasourceRecord.created_by.in_(actor_ids),
DatasourceRecord.updated_by.in_(actor_ids),
),
"datasource",
),
(
DatasourceMaterializationRecord,
DatasourceMaterializationRecord.created_by.in_(actor_ids),
"datasource_materialization",
),
(
DatasourcePayloadRecord,
DatasourcePayloadRecord.created_by.in_(actor_ids),
"datasource_payload",
),
(
DatasourceStageRecord,
DatasourceStageRecord.created_by.in_(actor_ids),
"datasource_stage",
),
(
DatasourcePublicationRecord,
DatasourcePublicationRecord.created_by.in_(actor_ids),
"datasource_publication",
),
(
DatasourceLifecycleEvidenceRecord,
DatasourceLifecycleEvidenceRecord.actor_ref.in_(actor_ids),
"datasource_lifecycle_evidence",
),
)
matches: list[_Match] = []
for model, condition, resource_type in specs:
rows = (
session.query(model)
.filter(model.tenant_id == tenant_id, condition)
.order_by(model.id)
.limit(_MAX_RECORDS + 1)
.all()
)
matches.extend(
_Match(resource_type, row, "datasource_operator_attribution")
for row in rows
)
if len(matches) > _MAX_RECORDS:
raise ValueError(
"Datasources DSAR result limit exceeded; narrow the selectors."
)
return matches
def _one(
session: Session,
model: Any,
*,
tenant_id: str,
field: str,
value: str,
) -> Any | None:
return (
session.query(model)
.filter(model.tenant_id == tenant_id, getattr(model, field) == value)
.one_or_none()
)
def _direct_category(session: Session, resource_type: str, row: Any) -> str:
if resource_type == "datasource_stage" and row.promoted_at is None:
return "unpromoted_datasource_stage"
if resource_type == "datasource_payload":
referenced = (
session.query(DatasourceMaterializationRecord.id)
.filter(
DatasourceMaterializationRecord.tenant_id == row.tenant_id,
DatasourceMaterializationRecord.payload_id == row.id,
)
.limit(1)
.count()
)
if not referenced:
return "unreferenced_datasource_payload"
if resource_type == "datasource_lifecycle_evidence":
return "datasource_operator_attribution"
return {
"datasource": "datasource_configuration",
"datasource_governance_reference": "datasource_governance_configuration",
"datasource_materialization": "immutable_datasource_materialization",
"datasource_payload": "referenced_datasource_payload",
"datasource_stage": "promoted_datasource_stage",
"datasource_publication": "immutable_datasource_publication",
"datasource_lifecycle_evidence": "datasource_operator_attribution",
}[resource_type]
def _root_datasource_ids(session: Session, match: _Match) -> set[str]:
row = match.row
if match.resource_type == "datasource":
return {row.id}
if match.resource_type == "datasource_stage":
return {row.target_datasource_id} if row.target_datasource_id else set()
if match.resource_type == "datasource_payload":
return {
value
for (value,) in session.query(DatasourceMaterializationRecord.datasource_id)
.filter(
DatasourceMaterializationRecord.tenant_id == row.tenant_id,
DatasourceMaterializationRecord.payload_id == row.id,
)
.all()
}
if match.resource_type == "datasource_lifecycle_evidence":
if row.subject_ref.startswith("datasource:"):
return {row.subject_ref.removeprefix("datasource:")}
if row.subject_ref.startswith("stage:"):
stage = session.get(
DatasourceStageRecord,
row.subject_ref.removeprefix("stage:"),
)
return {stage.target_datasource_id} if stage and stage.target_datasource_id else set()
if row.subject_ref.startswith("materialization:"):
materialization = session.get(
DatasourceMaterializationRecord,
row.subject_ref.removeprefix("materialization:"),
)
return {materialization.datasource_id} if materialization else set()
return set()
return {row.datasource_id}
def _correlates(match: _Match, actor_ids: tuple[str, ...]) -> bool:
row = match.row
return any(
str(getattr(row, field, "") or "") in actor_ids
for field in ("created_by", "updated_by", "actor_ref")
)
def _directly_targets(selectors: _Selectors, match: _Match) -> bool:
row = match.row
selector, field = {
"datasource": ("datasource_id", "id"),
"datasource_governance_reference": (
"governance_reference_id",
"id",
),
"datasource_materialization": ("materialization_id", "id"),
"datasource_payload": ("payload_id", "id"),
"datasource_stage": ("stage_id", "id"),
"datasource_publication": ("publication_id", "id"),
"datasource_lifecycle_evidence": ("lifecycle_evidence_id", "id"),
}[match.resource_type]
value = _strip_prefix(selectors.direct.get(selector, ""))
if value == str(getattr(row, field)):
return True
datasource_selector = _strip_prefix(selectors.direct.get("datasource_id", ""))
return bool(
datasource_selector
and datasource_selector in _root_datasource_ids_for_row(match)
)
def _root_datasource_ids_for_row(match: _Match) -> set[str]:
row = match.row
if match.resource_type == "datasource":
return {row.id}
if match.resource_type == "datasource_stage":
return {row.target_datasource_id} if row.target_datasource_id else set()
if match.resource_type == "datasource_payload":
return set()
if match.resource_type == "datasource_lifecycle_evidence":
return (
{row.subject_ref.removeprefix("datasource:")}
if row.subject_ref.startswith("datasource:")
else set()
)
return {row.datasource_id}
def _assert_deletable(session: Session, *, resource_type: str, row: Any) -> None:
if resource_type == "datasource_stage":
if row.promoted_at is not None or row.promoted_materialization_id is not None:
raise ValueError("Promoted Datasource stages require manual review.")
return
referenced = (
session.query(DatasourceMaterializationRecord.id)
.filter(
DatasourceMaterializationRecord.tenant_id == row.tenant_id,
DatasourceMaterializationRecord.payload_id == row.id,
)
.limit(1)
.count()
)
if referenced:
raise ValueError("Referenced Datasource payloads require manual review.")
def _record(match: _Match) -> DsarRecordRef:
row = match.row
immutable = match.category == "datasource_operator_attribution"
return DsarRecordRef(
provider_id="datasources",
module_id="datasources",
resource_type=match.resource_type,
resource_id=str(row.id),
category=match.category,
title=_title(match.resource_type),
data={
key: value
for key, value in _record_data(match.resource_type, row).items()
if value is not None
},
observed_at=_observed_at(row),
immutable_evidence=immutable,
retention_reason=(
"Institutional datasource creation, publication, and revision "
"activity remains attributable for governance and audit."
if immutable
else None
),
source_path="/datasources",
)
def _record_data(resource_type: str, row: Any) -> dict[str, object]:
if resource_type == "datasource":
return {
"kind": row.kind,
"mode": row.mode,
"shape": row.shape,
"status": row.status,
"schema_version": row.schema_version,
"row_count": row.row_count,
"byte_count": row.byte_count,
"authority_mode": row.authority_mode,
"classification": row.classification,
"publication_state": row.publication_state,
"deleted_at": _iso(row.deleted_at),
"created_at": _iso(row.created_at),
"updated_at": _iso(row.updated_at),
}
if resource_type == "datasource_lifecycle_evidence":
return {
"subject_ref": row.subject_ref,
"event_type": row.event_type,
"policy_version": row.policy_version,
"occurred_at": _iso(row.occurred_at),
}
if resource_type == "datasource_governance_reference":
return {"relation": row.relation}
if resource_type == "datasource_materialization":
return {
"revision": row.revision,
"state": row.state,
"schema_version": row.schema_version,
"row_count": row.row_count,
"byte_count": row.byte_count,
"frozen_at": _iso(row.frozen_at),
"source_timestamp": _iso(row.source_timestamp),
"created_at": _iso(row.created_at),
}
if resource_type == "datasource_payload":
return {
"backend": row.backend,
"state": row.state,
"media_type": row.media_type,
"row_count": row.row_count,
"byte_count": row.byte_count,
"created_at": _iso(row.created_at),
}
if resource_type == "datasource_stage":
return {
"kind": row.kind,
"mode": row.mode,
"shape": row.shape,
"state": row.state,
"row_count": row.row_count,
"byte_count": row.byte_count,
"promoted": row.promoted_at is not None,
"promoted_at": _iso(row.promoted_at),
"created_at": _iso(row.created_at),
}
return {
"producer_module": row.producer_module,
"status": row.status,
"created_at": _iso(row.created_at),
}
def _selectors(subject: DsarSubjectRef) -> _Selectors | None:
references = subject.external_references
account_id = _coalesce(
subject.account_id,
references.get("datasources.account"),
references.get("access.account"),
)
identity_id = _coalesce(
subject.identity_id,
references.get("datasources.identity"),
references.get("identity.id"),
)
membership_id = _coalesce(
subject.membership_id,
references.get("datasources.membership"),
references.get("tenancy.membership"),
)
if any(value is _CONFLICT for value in (account_id, identity_id, membership_id)):
return None
direct: dict[str, str] = {}
for selector, aliases in _DIRECT_ALIASES.items():
value = _coalesce(*(references.get(alias) for alias in aliases))
if value is _CONFLICT:
return None
if value:
direct[selector] = str(value)
return _Selectors(
account_id=_optional(account_id),
identity_id=_optional(identity_id),
membership_id=_optional(membership_id),
direct=direct,
)
def _coalesce(*values: str | None) -> str | None | object:
normalized = {str(value).strip() for value in values if str(value or "").strip()}
if len(normalized) > 1:
return _CONFLICT
return next(iter(normalized), None)
def _optional(value: object) -> str | None:
return value if isinstance(value, str) and value else None
def _strip_prefix(value: str) -> str:
return value.partition(":")[2] if ":" in value else value
def _title(resource_type: str) -> str:
return resource_type.replace("_", " ").title()
def _observed_at(row: Any) -> datetime | None:
for field in (
"promoted_at",
"source_timestamp",
"occurred_at",
"updated_at",
"created_at",
):
value = getattr(row, field, None)
if isinstance(value, datetime):
return _aware(value)
return None
def _iso(value: datetime | None) -> str | None:
aware = _aware(value)
return aware.isoformat() if aware else None
def _aware(value: datetime | None) -> datetime | None:
if value is None or value.tzinfo is not None:
return value
return value.replace(tzinfo=timezone.utc)
def _session(value: object) -> Session:
if not isinstance(value, Session):
raise TypeError("Datasources DSAR requires a SQLAlchemy Session.")
return value
def _validate_record(record: DsarRecordRef) -> None:
if record.provider_id != "datasources" or record.module_id != "datasources":
raise ValueError("Datasources DSAR cannot plan a foreign provider record.")
if record.resource_type not in _RESOURCE_MODELS or not record.resource_id:
raise ValueError("Datasources DSAR record identity is invalid.")
def _validate_action(action: DsarErasureActionRef) -> None:
if action.provider_id != "datasources" or action.module_id != "datasources":
raise ValueError("Datasources DSAR cannot execute a foreign provider action.")
if action.resource_type not in _RESOURCE_MODELS or not action.action_id.startswith(
"datasources:"
):
raise ValueError("Datasources DSAR action identity is invalid.")
__all__ = ["DATASOURCES_DSAR_CAPABILITY", "DatasourcesDsarProvider"]
@@ -0,0 +1,718 @@
from __future__ import annotations
import hashlib
import json
from collections.abc import Mapping, Sequence
from dataclasses import dataclass
from datetime import UTC, datetime, timedelta
from sqlalchemy import func, select
from sqlalchemy.orm import Session
from govoplan_core.core.datasources import DatasourceValidationError
from govoplan_core.db.base import utcnow
from govoplan_datasources.backend.db.models import (
DatasourceLifecycleEvidenceRecord,
DatasourceMaterializationRecord,
DatasourcePublicationRecord,
DatasourceRecord,
DatasourceStageRecord,
)
@dataclass(frozen=True, slots=True)
class RetentionCandidate:
ref: str
kind: str
datasource_ref: str | None
disposition: str
eligible_at: datetime
eligible: bool
blockers: tuple[str, ...]
policy_version: str
policy_hash: str
def to_dict(self) -> dict[str, object]:
return {
"ref": self.ref,
"kind": self.kind,
"datasource_ref": self.datasource_ref,
"disposition": self.disposition,
"eligible_at": _as_utc(self.eligible_at).isoformat(),
"eligible": self.eligible,
"blockers": list(self.blockers),
"policy_version": self.policy_version,
"policy_hash": self.policy_hash,
}
@dataclass(frozen=True, slots=True)
class RetentionPlan:
as_of: datetime
plan_hash: str
candidates: tuple[RetentionCandidate, ...]
def canonical_hash(value: object) -> str:
encoded = json.dumps(
value,
sort_keys=True,
separators=(",", ":"),
ensure_ascii=True,
default=str,
)
return hashlib.sha256(encoded.encode("utf-8")).hexdigest()
def normalize_approval_policy(value: Mapping[str, object] | None) -> dict[str, object]:
source = dict(value or {})
version = str(source.get("version") or "1").strip()
if not version:
raise DatasourceValidationError("Approval policy version is required.")
required = _boolean(source.get("required", False), label="required")
required_approvals = _bounded_integer(
source.get("required_approvals", 1),
label="required_approvals",
minimum=1,
maximum=5,
)
separation = _boolean(
source.get("separation_of_duties", True),
label="separation_of_duties",
)
expires_after_hours = source.get("expires_after_hours")
expiry = (
None
if expires_after_hours is None
else _bounded_integer(
expires_after_hours,
label="expires_after_hours",
minimum=1,
maximum=8_760,
)
)
policy_ref = _optional_text(source.get("policy_ref"))
return {
"version": version,
"required": required,
"required_approvals": required_approvals,
"separation_of_duties": separation,
"expires_after_hours": expiry,
"policy_ref": policy_ref,
}
def normalize_retention_policy(value: Mapping[str, object] | None) -> dict[str, object]:
source = dict(value or {})
version = str(source.get("version") or "1").strip()
if not version:
raise DatasourceValidationError("Retention policy version is required.")
enabled = _boolean(source.get("enabled", False), label="enabled")
durations = {
key: _optional_duration(source.get(key), label=key)
for key in (
"stage_days",
"materialization_days",
"frozen_evidence_days",
)
}
if enabled and all(value is None for value in durations.values()):
raise DatasourceValidationError(
"Enabled retention policy requires at least one retention duration."
)
return {
"version": version,
"enabled": enabled,
**durations,
"policy_ref": _optional_text(source.get("policy_ref")),
}
def initialize_stage_approval(
stage: DatasourceStageRecord,
*,
created_at: datetime | None = None,
) -> dict[str, object]:
policy = normalize_approval_policy(
_mapping(stage.governance_).get("approval_policy")
if isinstance(_mapping(stage.governance_).get("approval_policy"), Mapping)
else None
)
now = _as_utc(created_at or utcnow())
expiry_hours = policy["expires_after_hours"]
expires_at = (
now + timedelta(hours=int(expiry_hours))
if isinstance(expiry_hours, int)
else None
)
approval = {
"state": "pending" if policy["required"] else "not_required",
"policy": policy,
"policy_hash": canonical_hash(policy),
"subject_digest": stage_subject_digest(stage, policy=policy),
"required_approvals": policy["required_approvals"],
"approval_count": 0,
"expires_at": expires_at.isoformat() if expires_at else None,
"approvals": [],
}
stage.approval_ = approval
if stage.validation_.get("valid") is True:
stage.state = "awaiting_approval" if policy["required"] else "ready"
return approval
def decide_stage(
stage: DatasourceStageRecord,
*,
actor_ref: str,
actor_scopes: Sequence[str],
decision: str,
reason: str,
expected_policy_hash: str,
expected_subject_digest: str,
now: datetime | None = None,
) -> tuple[dict[str, object], bool]:
approval = dict(stage.approval_ or {})
if approval.get("policy_hash") != expected_policy_hash:
raise DatasourceValidationError(
"The approval policy changed; reload the stage before deciding."
)
if approval.get("subject_digest") != expected_subject_digest:
raise DatasourceValidationError(
"The staged content changed; reload the stage before deciding."
)
policy = normalize_approval_policy(_mapping(approval.get("policy")))
if not policy["required"]:
raise DatasourceValidationError("This stage does not require approval.")
if stage.state not in {"awaiting_approval", "ready"}:
raise DatasourceValidationError("This stage no longer accepts approval decisions.")
if policy["separation_of_duties"] and stage.created_by == actor_ref:
raise DatasourceValidationError(
"The stage creator cannot approve this stage under separation of duties."
)
decided_at = _as_utc(now or utcnow())
expires_at = _optional_datetime(approval.get("expires_at"))
if expires_at is not None and decided_at > expires_at:
approval["state"] = "expired"
stage.approval_ = approval
stage.state = "awaiting_approval"
raise DatasourceValidationError(
"The approval window expired; create a new stage for a fresh decision."
)
decisions = [dict(item) for item in _mapping_sequence(approval.get("approvals"))]
existing = next(
(item for item in decisions if item.get("actor_ref") == actor_ref),
None,
)
if existing is not None:
if existing.get("decision") == decision:
return approval, True
raise DatasourceValidationError(
"This approver already recorded a different decision."
)
cleaned_reason = reason.strip()
if not cleaned_reason:
raise DatasourceValidationError("An approval decision reason is required.")
decisions.append(
{
"actor_ref": actor_ref,
"decision": decision,
"reason": cleaned_reason,
"decided_at": decided_at.isoformat(),
"authority_scopes": sorted(set(actor_scopes)),
}
)
approved = sum(item.get("decision") == "approve" for item in decisions)
approval.update(
{
"approvals": decisions,
"approval_count": approved,
"state": (
"rejected"
if decision == "reject"
else (
"approved"
if approved >= int(policy["required_approvals"])
else "pending"
)
),
}
)
stage.approval_ = approval
stage.state = "rejected" if decision == "reject" else (
"ready"
if approval["state"] == "approved"
else "awaiting_approval"
)
return approval, False
def ensure_stage_approval_current(
stage: DatasourceStageRecord,
*,
current_policy: Mapping[str, object] | None,
) -> None:
approval = _mapping(stage.approval_)
stage_policy = normalize_approval_policy(_mapping(approval.get("policy")))
effective = normalize_approval_policy(current_policy)
if canonical_hash(effective) != approval.get("policy_hash"):
raise DatasourceValidationError(
"The approval policy changed after staging; create a new stage."
)
if stage_policy["required"] and approval.get("state") != "approved":
raise DatasourceValidationError("The stage has not reached its approval quorum.")
if approval.get("subject_digest") != stage_subject_digest(stage, policy=stage_policy):
raise DatasourceValidationError(
"The staged evidence changed after approval; create a new stage."
)
def stage_subject_digest(
stage: DatasourceStageRecord,
*,
policy: Mapping[str, object] | None = None,
) -> str:
return canonical_hash(
{
"stage_ref": _stage_ref(stage.id),
"target_datasource_ref": (
_datasource_ref(stage.target_datasource_id)
if stage.target_datasource_id
else None
),
"source_name": stage.source_name,
"mode": stage.mode,
"shape": stage.shape,
"fingerprint": stage.fingerprint,
"validation_policy_hash": stage.validation_.get("policy_hash"),
"approval_policy": dict(policy or {}),
}
)
def record_lifecycle_evidence(
session: Session,
*,
tenant_id: str,
subject_ref: str,
event_type: str,
actor_ref: str | None,
subject_digest: str,
details: Mapping[str, object],
policy_version: str | None = None,
policy_hash: str | None = None,
occurred_at: datetime | None = None,
) -> DatasourceLifecycleEvidenceRecord:
happened = _as_utc(occurred_at or utcnow())
previous_hash = session.scalar(
select(DatasourceLifecycleEvidenceRecord.event_hash)
.where(
DatasourceLifecycleEvidenceRecord.tenant_id == tenant_id,
DatasourceLifecycleEvidenceRecord.subject_ref == subject_ref,
)
.order_by(
DatasourceLifecycleEvidenceRecord.occurred_at.desc(),
DatasourceLifecycleEvidenceRecord.id.desc(),
)
.limit(1)
)
payload = {
"tenant_id": tenant_id,
"subject_ref": subject_ref,
"event_type": event_type,
"occurred_at": happened.isoformat(),
"actor_ref": actor_ref,
"policy_version": policy_version,
"policy_hash": policy_hash,
"subject_digest": subject_digest,
"previous_event_hash": previous_hash,
"details": dict(details),
}
item = DatasourceLifecycleEvidenceRecord(
tenant_id=tenant_id,
subject_ref=subject_ref,
event_type=event_type,
occurred_at=happened,
actor_ref=actor_ref,
policy_version=policy_version,
policy_hash=policy_hash,
subject_digest=subject_digest,
previous_event_hash=previous_hash,
event_hash=canonical_hash(payload),
details_=dict(details),
)
session.add(item)
session.flush()
return item
def list_lifecycle_evidence(
session: Session,
*,
tenant_id: str,
subject_ref: str | None = None,
limit: int = 200,
) -> tuple[DatasourceLifecycleEvidenceRecord, ...]:
statement = select(DatasourceLifecycleEvidenceRecord).where(
DatasourceLifecycleEvidenceRecord.tenant_id == tenant_id
)
if subject_ref:
statement = statement.where(
DatasourceLifecycleEvidenceRecord.subject_ref == subject_ref
)
statement = statement.order_by(
DatasourceLifecycleEvidenceRecord.occurred_at.desc(),
DatasourceLifecycleEvidenceRecord.id.desc(),
).limit(max(1, min(limit, 500)))
return tuple(session.scalars(statement))
def build_retention_plan(
session: Session,
*,
tenant_id: str,
as_of: datetime,
) -> RetentionPlan:
effective_at = _as_utc(as_of)
candidates: list[RetentionCandidate] = []
stages = session.scalars(
select(DatasourceStageRecord).where(
DatasourceStageRecord.tenant_id == tenant_id
)
)
for stage in stages:
policy = normalize_retention_policy(
_mapping(_mapping(stage.governance_).get("retention_policy"))
)
duration = policy.get("stage_days")
if not policy["enabled"] or not isinstance(duration, int):
continue
eligible_at = _as_utc(stage.created_at) + timedelta(days=duration)
if eligible_at > effective_at:
continue
blockers = (
("pending_approval",)
if stage.state == "awaiting_approval"
else ()
)
candidates.append(
RetentionCandidate(
ref=_stage_ref(stage.id),
kind="stage",
datasource_ref=(
_datasource_ref(stage.target_datasource_id)
if stage.target_datasource_id
else None
),
disposition="delete_stage",
eligible_at=eligible_at,
eligible=not blockers,
blockers=blockers,
policy_version=str(policy["version"]),
policy_hash=canonical_hash(policy),
)
)
materializations = session.execute(
select(DatasourceMaterializationRecord, DatasourceRecord)
.join(DatasourceRecord, DatasourceRecord.id == DatasourceMaterializationRecord.datasource_id)
.where(DatasourceMaterializationRecord.tenant_id == tenant_id)
)
for materialization, datasource in materializations:
if materialization.disposed_at is not None:
continue
snapshot = _mapping(materialization.governance_snapshot_)
policy = normalize_retention_policy(
_mapping(snapshot.get("retention_policy"))
)
duration_key = (
"frozen_evidence_days"
if materialization.frozen_at is not None
else "materialization_days"
)
duration = policy.get(duration_key)
if not policy["enabled"] or not isinstance(duration, int):
continue
eligible_at = _as_utc(materialization.created_at) + timedelta(days=duration)
if eligible_at > effective_at:
continue
blockers: list[str] = []
if datasource.current_materialization_id == materialization.id:
blockers.append("current_materialization")
if tuple(snapshot.get("hold_refs") or ()):
blockers.append("legal_hold")
publication_count = session.scalar(
select(func.count(DatasourcePublicationRecord.id)).where(
DatasourcePublicationRecord.materialization_id == materialization.id
)
)
if publication_count:
blockers.append("publication_evidence")
candidates.append(
RetentionCandidate(
ref=_materialization_ref(materialization.id),
kind="materialization",
datasource_ref=_datasource_ref(datasource.id),
disposition="purge_materialization_payload",
eligible_at=eligible_at,
eligible=not blockers,
blockers=tuple(blockers),
policy_version=str(policy["version"]),
policy_hash=canonical_hash(policy),
)
)
ordered = tuple(sorted(candidates, key=lambda item: (item.kind, item.ref)))
plan_payload = {
"tenant_id": tenant_id,
"as_of": effective_at.isoformat(),
"candidates": [item.to_dict() for item in ordered],
}
return RetentionPlan(
as_of=effective_at,
plan_hash=canonical_hash(plan_payload),
candidates=ordered,
)
def apply_retention_plan(
session: Session,
*,
tenant_id: str,
actor_ref: str,
plan: RetentionPlan,
target_refs: Sequence[str],
) -> tuple[tuple[str, ...], tuple[str, ...]]:
by_ref = {item.ref: item for item in plan.candidates}
requested = tuple(dict.fromkeys(target_refs))
missing = [ref for ref in requested if ref not in by_ref]
if missing:
raise DatasourceValidationError(
f"Retention targets are not in the current plan: {', '.join(missing)}."
)
blocked = [ref for ref in requested if not by_ref[ref].eligible]
if blocked:
raise DatasourceValidationError(
f"Retention targets have active blockers: {', '.join(blocked)}."
)
disposed: list[str] = []
evidence_hashes: list[str] = []
for ref in requested:
candidate = by_ref[ref]
if candidate.kind == "stage":
stage = _stage_by_ref(session, tenant_id=tenant_id, ref=ref)
digest = stage_subject_digest(
stage,
policy=normalize_approval_policy(
_mapping(_mapping(stage.approval_).get("policy"))
),
)
evidence = record_lifecycle_evidence(
session,
tenant_id=tenant_id,
subject_ref=ref,
event_type="retention.stage_deleted",
actor_ref=actor_ref,
subject_digest=digest,
policy_version=candidate.policy_version,
policy_hash=candidate.policy_hash,
details={
"plan_hash": plan.plan_hash,
"disposition": candidate.disposition,
"validation_policy_hash": stage.validation_.get("policy_hash"),
"approval_policy_hash": stage.approval_.get("policy_hash"),
},
occurred_at=plan.as_of,
)
evidence_hashes.append(evidence.event_hash)
session.delete(stage)
else:
materialization = _materialization_by_ref(
session,
tenant_id=tenant_id,
ref=ref,
)
payload = materialization.payload
subject_digest = canonical_hash(
{
"ref": ref,
"fingerprint": materialization.fingerprint,
"payload_checksum": materialization.payload_checksum,
"revision": materialization.revision,
}
)
evidence = record_lifecycle_evidence(
session,
tenant_id=tenant_id,
subject_ref=ref,
event_type="retention.materialization_payload_purged",
actor_ref=actor_ref,
subject_digest=subject_digest,
policy_version=candidate.policy_version,
policy_hash=candidate.policy_hash,
details={
"plan_hash": plan.plan_hash,
"disposition": candidate.disposition,
"payload_checksum": materialization.payload_checksum,
"row_count": materialization.row_count,
"byte_count": materialization.byte_count,
"frozen": materialization.frozen_at is not None,
},
occurred_at=plan.as_of,
)
evidence_hashes.append(evidence.event_hash)
materialization.payload_id = None
materialization.rows = []
materialization.state = "disposed"
materialization.disposed_at = plan.as_of
materialization.disposition_ = {
"plan_hash": plan.plan_hash,
"policy_version": candidate.policy_version,
"policy_hash": candidate.policy_hash,
"actor_ref": actor_ref,
"evidence_hash": evidence.event_hash,
"payload_checksum": materialization.payload_checksum,
}
session.flush()
if payload is not None:
remaining = session.scalar(
select(func.count(DatasourceMaterializationRecord.id)).where(
DatasourceMaterializationRecord.payload_id == payload.id
)
)
if not remaining:
session.delete(payload)
disposed.append(ref)
session.flush()
return tuple(disposed), tuple(evidence_hashes)
def _stage_by_ref(
session: Session,
*,
tenant_id: str,
ref: str,
) -> DatasourceStageRecord:
item = session.scalar(
select(DatasourceStageRecord).where(
DatasourceStageRecord.id == _ref_id(ref, "stage"),
DatasourceStageRecord.tenant_id == tenant_id,
)
)
if item is None:
raise DatasourceValidationError("Datasource stage is no longer available.")
return item
def _materialization_by_ref(
session: Session,
*,
tenant_id: str,
ref: str,
) -> DatasourceMaterializationRecord:
item = session.scalar(
select(DatasourceMaterializationRecord).where(
DatasourceMaterializationRecord.id == _ref_id(ref, "materialization"),
DatasourceMaterializationRecord.tenant_id == tenant_id,
)
)
if item is None:
raise DatasourceValidationError(
"Datasource materialization is no longer available."
)
return item
def _mapping(value: object) -> Mapping[str, object]:
return value if isinstance(value, Mapping) else {}
def _mapping_sequence(value: object) -> tuple[Mapping[str, object], ...]:
if not isinstance(value, Sequence) or isinstance(value, (str, bytes)):
return ()
return tuple(item for item in value if isinstance(item, Mapping))
def _boolean(value: object, *, label: str) -> bool:
if not isinstance(value, bool):
raise DatasourceValidationError(f"Approval or retention {label} must be boolean.")
return value
def _bounded_integer(
value: object,
*,
label: str,
minimum: int,
maximum: int,
) -> int:
if isinstance(value, bool) or not isinstance(value, int):
raise DatasourceValidationError(f"Policy {label} must be an integer.")
if value < minimum or value > maximum:
raise DatasourceValidationError(
f"Policy {label} must be between {minimum} and {maximum}."
)
return value
def _optional_duration(value: object, *, label: str) -> int | None:
if value is None:
return None
return _bounded_integer(value, label=label, minimum=1, maximum=36_500)
def _optional_text(value: object) -> str | None:
text = str(value or "").strip()
return text or None
def _optional_datetime(value: object) -> datetime | None:
if not value:
return None
if isinstance(value, datetime):
return _as_utc(value)
try:
return _as_utc(datetime.fromisoformat(str(value)))
except ValueError as exc:
raise DatasourceValidationError("Approval expiry is invalid.") from exc
def _as_utc(value: datetime) -> datetime:
if value.tzinfo is None:
return value.replace(tzinfo=UTC)
return value.astimezone(UTC)
def _ref_id(ref: str, prefix: str) -> str:
expected = f"{prefix}:"
if not ref.startswith(expected) or not ref.removeprefix(expected).strip():
raise DatasourceValidationError(f"Invalid {prefix} reference.")
return ref.removeprefix(expected)
def _stage_ref(value: str) -> str:
return f"stage:{value}"
def _datasource_ref(value: str) -> str:
return f"datasource:{value}"
def _materialization_ref(value: str) -> str:
return f"materialization:{value}"
__all__ = [
"RetentionCandidate",
"RetentionPlan",
"apply_retention_plan",
"build_retention_plan",
"canonical_hash",
"decide_stage",
"ensure_stage_approval_current",
"initialize_stage_approval",
"list_lifecycle_evidence",
"normalize_approval_policy",
"normalize_retention_policy",
"record_lifecycle_evidence",
"stage_subject_digest",
]
+309 -21
View File
@@ -7,16 +7,21 @@ from govoplan_core.core.access import (
CAPABILITY_AUTH_PRINCIPAL_RESOLVER, CAPABILITY_AUTH_PRINCIPAL_RESOLVER,
) )
from govoplan_core.core.datasources import ( from govoplan_core.core.datasources import (
CAPABILITY_DATASOURCE_ARTIFACT_BACKENDS,
CAPABILITY_DATASOURCE_CATALOGUE, CAPABILITY_DATASOURCE_CATALOGUE,
CAPABILITY_DATASOURCE_LIFECYCLE, CAPABILITY_DATASOURCE_LIFECYCLE,
CAPABILITY_DATASOURCE_ORIGINS, CAPABILITY_DATASOURCE_ORIGINS,
CAPABILITY_DATASOURCE_PUBLICATION, CAPABILITY_DATASOURCE_PUBLICATION,
CAPABILITY_POLICY_DATASOURCE_VISIBILITY,
datasource_artifact_backend_provider,
) )
from govoplan_core.core.module_guards import ( from govoplan_core.core.module_guards import (
drop_table_retirement_provider, drop_table_retirement_provider,
persistent_table_uninstall_guard, persistent_table_uninstall_guard,
) )
from govoplan_core.core.modules import ( from govoplan_core.core.modules import (
CapabilityDocumentation,
DocumentationCondition,
DocumentationLink, DocumentationLink,
DocumentationTopic, DocumentationTopic,
FrontendModule, FrontendModule,
@@ -28,6 +33,7 @@ from govoplan_core.core.modules import (
ModuleManifest, ModuleManifest,
NavItem, NavItem,
PermissionDefinition, PermissionDefinition,
ProductAreaContribution,
RoleTemplate, RoleTemplate,
) )
from govoplan_core.core.provider_governance import ( from govoplan_core.core.provider_governance import (
@@ -35,22 +41,32 @@ from govoplan_core.core.provider_governance import (
ModuleArchitectureDocumentation, ModuleArchitectureDocumentation,
ModuleMaturityEvidence, ModuleMaturityEvidence,
) )
from govoplan_core.core.search import SearchSourceProviderRegistration
from govoplan_core.core.views import ViewSurface from govoplan_core.core.views import ViewSurface
from govoplan_core.db.base import Base from govoplan_core.db.base import Base
from govoplan_datasources.backend.db import models as datasource_models from govoplan_datasources.backend.db import models as datasource_models
from govoplan_datasources.backend.dsar_provider import (
DATASOURCES_DSAR_CAPABILITY,
DatasourcesDsarProvider,
)
from govoplan_datasources.backend.search_source import (
create_datasources_search_source,
)
from govoplan_datasources.backend.service import ( from govoplan_datasources.backend.service import (
ADMIN_SCOPE, ADMIN_SCOPE,
CATALOGUE_READ_SCOPE, CATALOGUE_READ_SCOPE,
SOURCE_WRITE_SCOPE, SOURCE_WRITE_SCOPE,
STAGE_APPROVE_SCOPE,
STAGE_WRITE_SCOPE, STAGE_WRITE_SCOPE,
SqlDatasourceProvider, SqlDatasourceProvider,
) )
from govoplan_datasources.backend.payloads import ExternalArtifactPayloadBackend
MODULE_ID = "datasources" MODULE_ID = "datasources"
MODULE_NAME = "Datasources" MODULE_NAME = "Datasources"
MODULE_VERSION = "0.1.18" MODULE_VERSION = "0.1.21"
DATASOURCE_INTERFACE_VERSION = "0.1.0" DATASOURCE_INTERFACE_VERSION = "0.2.0"
ARCHITECTURE = ModuleArchitectureDeclaration( ARCHITECTURE = ModuleArchitectureDeclaration(
layer="data_reporting_integration", layer="data_reporting_integration",
@@ -75,7 +91,7 @@ ARCHITECTURE = ModuleArchitectureDeclaration(
), ),
known_limits=( known_limits=(
"Governance references are stable provider-neutral refs; dedicated selectors depend on the owning optional modules.", "Governance references are stable provider-neutral refs; dedicated selectors depend on the owning optional modules.",
"Quality and freshness policies are stored and snapshotted but enforcement remains provider-specific.", "External artifact content validation depends on an installed payload backend and checksum-bound producer evidence.",
), ),
supported_authority_modes=( supported_authority_modes=(
"native_authoritative", "native_authoritative",
@@ -135,6 +151,11 @@ PERMISSIONS = (
"Stage datasource content", "Stage datasource content",
"Upload, validate, inspect, and promote bounded datasource stages.", "Upload, validate, inspect, and promote bounded datasource stages.",
), ),
_permission(
STAGE_APPROVE_SCOPE,
"Approve datasource promotion",
"Review staged validation evidence and record an attributable promotion decision.",
),
_permission( _permission(
ADMIN_SCOPE, ADMIN_SCOPE,
"Administer datasources", "Administer datasources",
@@ -153,6 +174,12 @@ ROLE_TEMPLATES = (
STAGE_WRITE_SCOPE, STAGE_WRITE_SCOPE,
), ),
), ),
RoleTemplate(
slug="datasource_approver",
name="Datasource approver",
description="Independently approve or reject governed datasource stages.",
permissions=(CATALOGUE_READ_SCOPE, STAGE_APPROVE_SCOPE),
),
RoleTemplate( RoleTemplate(
slug="datasource_reader", slug="datasource_reader",
name="Datasource reader", name="Datasource reader",
@@ -172,7 +199,24 @@ def _router(context: ModuleContext):
def _provider(context: ModuleContext) -> SqlDatasourceProvider: def _provider(context: ModuleContext) -> SqlDatasourceProvider:
return SqlDatasourceProvider(registry=context.registry) provider = datasource_artifact_backend_provider(context.registry)
backends = (
tuple(
ExternalArtifactPayloadBackend(backend)
for backend in provider.artifact_backends()
)
if provider is not None
else ()
)
return SqlDatasourceProvider(
registry=context.registry,
payload_backends=backends,
)
def _dsar_provider(context: ModuleContext) -> DatasourcesDsarProvider:
del context
return DatasourcesDsarProvider()
def _tenant_summary(session, tenant_id: str) -> dict[str, int]: def _tenant_summary(session, tenant_id: str) -> dict[str, int]:
@@ -204,6 +248,15 @@ def _tenant_summary(session, tenant_id: str) -> dict[str, int]:
) )
.count() .count()
), ),
"datasource_stages_awaiting_approval": (
session.query(datasource_models.DatasourceStageRecord)
.filter(
datasource_models.DatasourceStageRecord.tenant_id == tenant_id,
datasource_models.DatasourceStageRecord.state
== "awaiting_approval",
)
.count()
),
} }
@@ -219,11 +272,14 @@ manifest = ModuleManifest(
"files", "files",
"notifications", "notifications",
"policy", "policy",
"search",
), ),
optional_capabilities=( optional_capabilities=(
CAPABILITY_AUTH_PRINCIPAL_RESOLVER, CAPABILITY_AUTH_PRINCIPAL_RESOLVER,
CAPABILITY_AUTH_PERMISSION_EVALUATOR, CAPABILITY_AUTH_PERMISSION_EVALUATOR,
CAPABILITY_DATASOURCE_ORIGINS, CAPABILITY_DATASOURCE_ORIGINS,
CAPABILITY_DATASOURCE_ARTIFACT_BACKENDS,
CAPABILITY_POLICY_DATASOURCE_VISIBILITY,
), ),
provides_interfaces=( provides_interfaces=(
ModuleInterfaceProvider( ModuleInterfaceProvider(
@@ -246,6 +302,7 @@ manifest = ModuleManifest(
name="datasources.publication", name="datasources.publication",
version=DATASOURCE_INTERFACE_VERSION, version=DATASOURCE_INTERFACE_VERSION,
), ),
ModuleInterfaceProvider(name=DATASOURCES_DSAR_CAPABILITY, version="0.1.0"),
), ),
requires_interfaces=( requires_interfaces=(
ModuleInterfaceRequirement( ModuleInterfaceRequirement(
@@ -254,6 +311,18 @@ manifest = ModuleManifest(
version_max_exclusive="1.0.0", version_max_exclusive="1.0.0",
optional=True, optional=True,
), ),
ModuleInterfaceRequirement(
name="search.source",
version_min="1.0.0",
version_max_exclusive="2.0.0",
optional=True,
),
ModuleInterfaceRequirement(
name="policy.datasource_visibility",
version_min="1.0.0",
version_max_exclusive="2.0.0",
optional=True,
),
), ),
permissions=PERMISSIONS, permissions=PERMISSIONS,
role_templates=ROLE_TEMPLATES, role_templates=ROLE_TEMPLATES,
@@ -286,13 +355,77 @@ manifest = ModuleManifest(
order=70, order=70,
), ),
), ),
product_areas=(
ProductAreaContribution(
id="data-assurance",
module_id=MODULE_ID,
label="i18n:govoplan-core.product_area.data_assurance",
icon="database-zap",
description="i18n:govoplan-core.product_area.data_assurance_description",
surface_ids=(
"datasources.nav.datasources",
"datasources.route.datasources",
),
order=60,
),
),
view_surfaces=( view_surfaces=(
ViewSurface(id="datasources.page", module_id=MODULE_ID, kind="route", label="Datasources", order=70), ViewSurface(
ViewSurface(id="datasources.catalogue", module_id=MODULE_ID, kind="section", label="Datasource catalogue", order=10), id="datasources.page",
ViewSurface(id="datasources.staging", module_id=MODULE_ID, kind="section", label="Datasource staging", order=20), module_id=MODULE_ID,
ViewSurface(id="datasources.origins", module_id=MODULE_ID, kind="section", label="Datasource origins", order=30), kind="route",
ViewSurface(id="datasources.governance", module_id=MODULE_ID, kind="action", label="Datasource governance", order=40), label="Datasources",
ViewSurface(id="datasources.preview", module_id=MODULE_ID, kind="section", label="Datasource preview and materializations", order=50), order=70,
),
ViewSurface(
id="datasources.catalogue",
module_id=MODULE_ID,
kind="section",
label="Datasource catalogue",
order=10,
),
ViewSurface(
id="datasources.staging",
module_id=MODULE_ID,
kind="section",
label="Datasource staging",
order=20,
),
ViewSurface(
id="datasources.origins",
module_id=MODULE_ID,
kind="section",
label="Datasource origins",
order=30,
),
ViewSurface(
id="datasources.governance",
module_id=MODULE_ID,
kind="action",
label="Datasource governance",
order=40,
),
ViewSurface(
id="datasources.preview",
module_id=MODULE_ID,
kind="section",
label="Datasource preview and materializations",
order=50,
),
ViewSurface(
id="datasources.lifecycle-evidence",
module_id=MODULE_ID,
kind="section",
label="Datasource lifecycle evidence",
order=60,
),
ViewSurface(
id="datasources.retention",
module_id=MODULE_ID,
kind="action",
label="Datasource retention preview and apply",
order=70,
),
), ),
), ),
route_factory=_router, route_factory=_router,
@@ -300,14 +433,31 @@ manifest = ModuleManifest(
CAPABILITY_DATASOURCE_CATALOGUE: _provider, CAPABILITY_DATASOURCE_CATALOGUE: _provider,
CAPABILITY_DATASOURCE_LIFECYCLE: _provider, CAPABILITY_DATASOURCE_LIFECYCLE: _provider,
CAPABILITY_DATASOURCE_PUBLICATION: _provider, CAPABILITY_DATASOURCE_PUBLICATION: _provider,
DATASOURCES_DSAR_CAPABILITY: _dsar_provider,
},
capability_documentation={
DATASOURCES_DSAR_CAPABILITY: CapabilityDocumentation(
label="Datasources data-subject request provider",
summary="Finds governed datasource copies without exposing credentials or row payloads.",
contract_version="0.1.0",
documentation_types=("admin", "user"),
audience=("privacy_officer", "data_steward", "user"),
),
}, },
tenant_summary_providers=(_tenant_summary,), tenant_summary_providers=(_tenant_summary,),
search_sources=(
SearchSourceProviderRegistration(
id="datasources.catalogue",
factory=create_datasources_search_source,
),
),
migration_spec=MigrationSpec( migration_spec=MigrationSpec(
module_id=MODULE_ID, module_id=MODULE_ID,
metadata=Base.metadata, metadata=Base.metadata,
script_location=str(Path(__file__).with_name("migrations") / "versions"), script_location=str(Path(__file__).with_name("migrations") / "versions"),
retirement_supported=True, retirement_supported=True,
retirement_provider=drop_table_retirement_provider( retirement_provider=drop_table_retirement_provider(
datasource_models.DatasourceLifecycleEvidenceRecord,
datasource_models.DatasourcePublicationRecord, datasource_models.DatasourcePublicationRecord,
datasource_models.DatasourceStageRecord, datasource_models.DatasourceStageRecord,
datasource_models.DatasourceMaterializationRecord, datasource_models.DatasourceMaterializationRecord,
@@ -323,6 +473,7 @@ manifest = ModuleManifest(
), ),
uninstall_guard_providers=( uninstall_guard_providers=(
persistent_table_uninstall_guard( persistent_table_uninstall_guard(
datasource_models.DatasourceLifecycleEvidenceRecord,
datasource_models.DatasourceRecord, datasource_models.DatasourceRecord,
datasource_models.DatasourceMaterializationRecord, datasource_models.DatasourceMaterializationRecord,
datasource_models.DatasourcePayloadRecord, datasource_models.DatasourcePayloadRecord,
@@ -334,6 +485,29 @@ manifest = ModuleManifest(
), ),
architecture=ARCHITECTURE, architecture=ARCHITECTURE,
documentation=( documentation=(
DocumentationTopic(
id="datasources.data-subject-requests",
title="Datasources data-subject requests",
summary="Identify governed datasource copies while preserving credential, payload, and immutable-evidence boundaries.",
body=(
"Datasources matches exact tenant-scoped catalogue, governance-reference, materialization, payload, stage, publication, and lifecycle-evidence identifiers plus minimized account, identity, or membership operator attribution. DSAR results never copy connector/provider references, locators, credentials, arbitrary rows, schemas, validation samples, metadata, provenance bodies, checkpoints, idempotency material, or hashes. Arbitrary tabular payloads are not scanned for identifiers because that would be incomplete, schema-dependent, and liable to disclose unrelated people; the authoritative source module must locate and correct subject facts. "
"An unpromoted stage or unreferenced payload can be deleted idempotently. Published catalogue state, promoted stages, referenced payloads, immutable materializations, governance references, publications, holds, and operator attribution require data-steward review and source correction. Downstream Dataflow and Reporting outputs must be refreshed after correction."
),
layer="configured",
documentation_types=("admin", "user"),
audience=("user", "operator", "module_admin", "data_steward", "auditor"),
related_modules=("core", "connectors", "dataflow", "reporting"),
order=69,
metadata={
"seed": True,
"help_contexts": [
"datasources.data-subject-requests",
"datasources.catalogue",
"datasources.staging",
"datasources.preview",
],
},
),
DocumentationTopic( DocumentationTopic(
id="datasources.lifecycle", id="datasources.lifecycle",
title="Datasource lifecycle", title="Datasource lifecycle",
@@ -350,7 +524,11 @@ manifest = ModuleManifest(
"reports, controls, and decisions. It preserves origin source mode, " "reports, controls, and decisions. It preserves origin source mode, "
"structured health, declared pushdown, and effective row, byte, and time " "structured health, declared pushdown, and effective row, byte, and time "
"limits for live previews. Governance metadata visibility does not " "limits for live previews. Governance metadata visibility does not "
"grant access to protected rows." "grant access to protected rows. When Search is enabled, the module "
"indexes only bounded catalogue labels and governance-safe facets, "
"then rechecks the current catalogue permission before returning a "
"result. Rows, schemas, connector references, credentials, arbitrary "
"metadata, and provenance are never copied into the search index."
), ),
layer="available", layer="available",
documentation_types=("admin", "user"), documentation_types=("admin", "user"),
@@ -362,6 +540,7 @@ manifest = ModuleManifest(
"files", "files",
"reporting", "reporting",
"risk_compliance", "risk_compliance",
"search",
), ),
order=70, order=70,
metadata={ metadata={
@@ -383,7 +562,10 @@ manifest = ModuleManifest(
"Authority mode states whether GovOPlaN, an external system, a synchronized projection, an overlay, or a linked reference " "Authority mode states whether GovOPlaN, an external system, a synchronized projection, an overlay, or a linked reference "
"controls the data. The authoritative source, owner, steward, responsible organization/function, schema owner, privacy " "controls the data. The authoritative source, owner, steward, responsible organization/function, schema owner, privacy "
"profile, retention policy, transfer agreement, legal basis, holds, correction procedure, purposes, official keys, and " "profile, retention policy, transfer agreement, legal basis, holds, correction procedure, purposes, official keys, and "
"known limits provide discoverable institutional context. Freshness and quality policies are typed JSON contracts retained " "known limits provide discoverable institutional context. A retention-policy reference identifies an owning external Policy rule, while "
"the local versioned retention contract controls previewable stage and payload disposition. Retention never overrides legal holds, the current materialization, or publication evidence. A transfer-agreement "
"reference records the governed agreement for external exchange; it does not grant connector credentials, recipient access, "
"or authority to export classified rows. Freshness and quality policies are typed JSON contracts retained "
"with materialization evidence. Datasources enforces the declared bounded tabular stage rules and schema policy; origin-specific " "with materialization evidence. Datasources enforces the declared bounded tabular stage rules and schema policy; origin-specific "
"or consumer-specific controls still remain with the provider or consuming control that declares support. Metadata " "or consumer-specific controls still remain with the provider or consuming control that declares support. Metadata "
"visibility never grants row access." "visibility never grants row access."
@@ -391,7 +573,14 @@ manifest = ModuleManifest(
layer="available", layer="available",
documentation_types=("admin", "user"), documentation_types=("admin", "user"),
audience=("operator", "module_admin", "data_steward", "product_owner"), audience=("operator", "module_admin", "data_steward", "product_owner"),
related_modules=("policy", "organizations", "idm", "dataflow", "reporting", "risk_compliance"), related_modules=(
"policy",
"organizations",
"idm",
"dataflow",
"reporting",
"risk_compliance",
),
order=71, order=71,
metadata={ metadata={
"seed": True, "seed": True,
@@ -403,6 +592,38 @@ manifest = ModuleManifest(
"datasources.field.publication-state", "datasources.field.publication-state",
"datasources.field.freshness-policy", "datasources.field.freshness-policy",
"datasources.field.quality-policy", "datasources.field.quality-policy",
"datasources.field.retention-policy",
"datasources.field.access-policy",
"datasources.field.visibility-policy",
"datasources.field.transfer-agreement",
],
},
),
DocumentationTopic(
id="datasources.visibility",
title="Datasource and field visibility",
summary="Apply source, materialization, field, and row controls before protected rows leave Datasources.",
body=(
"A local visibility policy can restrict discovery and reads by account, membership, identity, group, role, service account, or authentication method. "
"Materialization ACLs additionally protect current and frozen revisions. Field rules classify a field and either replace its value with null or omit the field unless an allow ACL matches. Row filters compare a configured field with a principal account, membership, identity, service-account, group, role, or function-assignment claim; all configured filters and all policy overlays must allow a row. "
"Datasources applies these controls before rows leave its provider boundary and returns an opaque permitted-view fingerprint bound to the base revision, effective policies, and relevant principal facts. Denied and filtered reads emit audit evidence containing outcomes and policy diagnostics, never protected values, provider locators, arbitrary metadata, or provenance bodies. "
"An optional access-policy reference delegates hierarchical overlays to Policy. Policy can only tighten local behavior. An unresolved reference, unavailable configured provider, or malformed policy fails closed. Without Policy and without a reference, the local visibility policy remains fully usable. Frozen reads enforce the current local policy together with a distinct restrictive policy captured in the frozen governance snapshot."
),
layer="available",
documentation_types=("admin", "user"),
audience=("operator", "module_admin", "data_steward", "auditor"),
related_modules=("policy", "access", "audit", "connectors"),
order=72,
metadata={
"seed": True,
"help_contexts": [
"datasources.field.access-policy",
"datasources.field.visibility-policy",
"datasources.preview",
],
"limitations": [
"Row filters support deterministic equality or membership against principal claims; arbitrary expressions are not executed.",
"Live filtered previews scan at most 10,000 origin rows and disclose a structured limitation diagnostic if the origin is larger.",
], ],
}, },
), ),
@@ -416,8 +637,11 @@ manifest = ModuleManifest(
"promotion. Updates compare the detected schema with the current target and classify each change as compatible, warning, or " "promotion. Updates compare the detected schema with the current target and classify each change as compatible, warning, or "
"breaking. Diagnostics expose counts and bounded row numbers, never field values. The policy version and hash, diagnostics, and " "breaking. Diagnostics expose counts and bounded row numbers, never field values. The policy version and hash, diagnostics, and "
"schema diff are copied into immutable materialization provenance when promotion succeeds. Producer publication applies the same " "schema diff are copied into immutable materialization provenance when promotion succeeds. Producer publication applies the same "
"gate before any catalogue effect, retains validation evidence on the immutable output revision, serializes a publication identity across PostgreSQL worker nodes before replay lookup, and emits a transactional terminal event. Approval and retention execution " "gate before any catalogue effect, retains validation evidence on the immutable output revision, serializes a publication identity across PostgreSQL worker nodes before replay lookup, and emits a transactional terminal event. Large outputs may instead supply a "
"are not inferred from arbitrary JSON flags and remain separate governed lifecycle work." "durable provider-neutral artifact reference with a pinned locator, SHA-256 checksum, declared schema, fingerprint, and size. "
"The configured payload backend verifies the artifact and provides bounded reads. Content-level rules require checksum- and "
"policy-bound producer evidence; absent evidence creates an immutable review-required materialization without changing the "
"Datasource's current state. Warning and review-required outcomes are preserved for Workflow handoffs. Versioned approval policy can require an attributable, separated quorum before a stage or refresh becomes current; retention uses a fresh, hashed preview before any explicit administrator apply."
), ),
layer="available", layer="available",
documentation_types=("admin", "user"), documentation_types=("admin", "user"),
@@ -429,8 +653,14 @@ manifest = ModuleManifest(
kind="repository", kind="repository",
), ),
), ),
related_modules=("policy", "approvals", "audit", "dataflow", "workflow_engine"), related_modules=(
order=72, "policy",
"approvals",
"audit",
"dataflow",
"workflow_engine",
),
order=73,
metadata={ metadata={
"seed": True, "seed": True,
"help_contexts": [ "help_contexts": [
@@ -441,7 +671,9 @@ manifest = ModuleManifest(
], ],
"limitations": [ "limitations": [
"Referential rules currently use a bounded embedded value set rather than reading another protected Datasource.", "Referential rules currently use a bounded embedded value set rather than reading another protected Datasource.",
"Approval authority and automatic retention execution are not part of the current stage contract.", "Approvals are module-local lifecycle evidence; installing a separate Approvals module does not silently change configured datasource policy.",
"Retention is never automatic: an administrator must preview a current plan and explicitly apply eligible targets.",
"Artifact bytes remain owned by their payload backend; Datasources stores an immutable reference and integrity evidence.",
], ],
}, },
), ),
@@ -453,7 +685,7 @@ manifest = ModuleManifest(
"A Datasource key is the stable catalogue identity used by consumers. Live mode reads through an available origin; cached " "A Datasource key is the stable catalogue identity used by consumers. Live mode reads through an available origin; cached "
"mode refreshes an origin into immutable revisions; static mode promotes uploaded content from staging. Stages are bounded, " "mode refreshes an origin into immutable revisions; static mode promotes uploaded content from staging. Stages are bounded, "
"inspectable, and non-consumable until promoted. Promotion creates or updates a governed Datasource and appends an immutable " "inspectable, and non-consumable until promoted. Promotion creates or updates a governed Datasource and appends an immutable "
"materialization. Refresh appends a new cached revision without rewriting older evidence. Freeze labels an immutable, " "materialization. When approval is configured, a refresh first creates a non-consumable stage and only an approved quorum permits promotion. Freeze labels an immutable, "
"addressable state for reproducible execution. Retirement removes the Datasource from new definitions while retained " "addressable state for reproducible execution. Retirement removes the Datasource from new definitions while retained "
"materialization references remain governed. Connector absence disables origin registration but leaves local catalogue and " "materialization references remain governed. Connector absence disables origin registration but leaves local catalogue and "
"staging behavior available." "staging behavior available."
@@ -461,8 +693,14 @@ manifest = ModuleManifest(
layer="available", layer="available",
documentation_types=("admin", "user"), documentation_types=("admin", "user"),
audience=("operator", "module_admin", "power_user", "product_owner"), audience=("operator", "module_admin", "power_user", "product_owner"),
related_modules=("connectors", "dataflow", "workflow_engine", "reporting", "audit"), related_modules=(
order=73, "connectors",
"dataflow",
"workflow_engine",
"reporting",
"audit",
),
order=74,
metadata={ metadata={
"seed": True, "seed": True,
"help_contexts": [ "help_contexts": [
@@ -482,6 +720,56 @@ manifest = ModuleManifest(
}, },
}, },
), ),
DocumentationTopic(
id="datasources.approval-and-retention",
title="Approve promotions and apply retention safely",
summary="Require separated datasource decisions and preview retention consequences before disposing staged or materialized payloads.",
body=(
"A versioned approval policy can require one to five distinct approvals before a valid stage becomes ready. The stage creator cannot approve when separation of duties is enabled. Decisions are bound to the exact stage fingerprint, quality-policy hash, approval-policy hash, actor authority, reason, and expiry; a changed policy or staged subject invalidates promotion. Cached refreshes use the same stage path whenever approval is required. "
"A versioned retention policy independently sets durations for stages, ordinary materializations, and frozen evidence. Preview returns a hashed plan with every due target and blocker. Current materializations, legal holds, pending approvals, and producer-publication evidence cannot be disposed. Applying the unchanged plan deletes eligible transient stages or purges materialization payload rows while retaining minimized schema/provenance, checksum, disposition, and hash-chained lifecycle evidence. Empty local policies disable both gates, so the module remains usable without Access or Policy; installed modules may tighten access but do not manufacture approval claims."
),
layer="available",
documentation_types=("admin", "user"),
audience=("operator", "module_admin", "data_steward", "auditor"),
conditions=(
DocumentationCondition(
any_scopes=(
STAGE_WRITE_SCOPE,
STAGE_APPROVE_SCOPE,
ADMIN_SCOPE,
)
),
),
related_modules=("access", "audit", "policy", "workflow_engine"),
order=75,
translations={
"de": {
"title": "Freigaben erteilen und Aufbewahrung sicher anwenden",
"summary": "Getrennte Entscheidungen für Datenquellen verlangen und Folgen der Aufbewahrung vor dem Löschen von Stufen oder Nutzdaten prüfen.",
"body": (
"Eine versionierte Freigaberichtlinie kann ein bis fünf verschiedene Freigaben verlangen, bevor eine gültige Stufe bereit ist. Bei aktivierter Funktionstrennung darf die erstellende Person nicht selbst freigeben. Entscheidungen sind an Fingerabdruck, Qualitäts- und Freigaberichtlinien-Hash, Berechtigung, Begründung und Ablaufzeit gebunden; geänderte Richtlinien oder Inhalte verhindern die Übernahme. Zwischengespeicherte Aktualisierungen nutzen bei Freigabepflicht denselben Stufenweg. "
"Eine getrennte Aufbewahrungsrichtlinie legt Fristen für Stufen, gewöhnliche Materialisierungen und eingefrorene Nachweise fest. Die Vorschau liefert einen gehashten Plan mit fälligen Zielen und Sperrgründen. Aktuelle Materialisierungen, rechtliche Sperren, offene Freigaben und Veröffentlichungsnachweise dürfen nicht entfernt werden. Beim Anwenden des unveränderten Plans werden nur zulässige Stufen oder Nutzdaten gelöscht; minimierte Metadaten, Prüfsummen, Dispositionsangaben und hashverkettete Lebenszyklusnachweise bleiben erhalten. Leere lokale Richtlinien deaktivieren beide Schranken, sodass das Modul ohne Access oder Policy nutzbar bleibt."
),
}
},
metadata={
"seed": True,
"kind": "workflow",
"help_contexts": [
"datasources.staging.approval",
"datasources.action.approve",
"datasources.action.reject",
"datasources.field.approval-policy",
"datasources.field.retention-policy-contract",
"datasources.lifecycle-evidence",
"datasources.retention",
],
"limitations": [
"Retention execution is explicit and plan-bound; this module does not run a hidden deletion scheduler.",
"Disposed materialization records retain minimized schema and provenance metadata while their payload rows are removed.",
],
},
),
), ),
) )
@@ -0,0 +1,38 @@
"""v0.1.20 datasource visibility policy
Revision ID: c9e3a6f1d4b8
Revises: b8d2f5a0c3e7
Create Date: 2026-08-21 20:15:00.000000
"""
from __future__ import annotations
from alembic import op
import sqlalchemy as sa
revision = "c9e3a6f1d4b8"
down_revision = "b8d2f5a0c3e7"
branch_labels = None
depends_on = None
def upgrade() -> None:
op.add_column(
"datasource_catalogue",
sa.Column("access_policy_ref", sa.String(length=500), nullable=True),
)
op.add_column(
"datasource_catalogue",
sa.Column(
"visibility_policy",
sa.JSON(),
server_default=sa.text("'{}'"),
nullable=False,
),
)
def downgrade() -> None:
op.drop_column("datasource_catalogue", "visibility_policy")
op.drop_column("datasource_catalogue", "access_policy_ref")
@@ -0,0 +1,142 @@
"""v0.1.21 datasource lifecycle governance
Revision ID: d1a7c3e9f5b2
Revises: c9e3a6f1d4b8
Create Date: 2026-08-22 20:00:00.000000
"""
from __future__ import annotations
from alembic import op
import sqlalchemy as sa
revision = "d1a7c3e9f5b2"
down_revision = "c9e3a6f1d4b8"
branch_labels = None
depends_on = None
def upgrade() -> None:
op.add_column(
"datasource_catalogue",
sa.Column(
"approval_policy",
sa.JSON(),
server_default=sa.text("'{}'"),
nullable=False,
),
)
op.add_column(
"datasource_catalogue",
sa.Column(
"retention_policy",
sa.JSON(),
server_default=sa.text("'{}'"),
nullable=False,
),
)
op.add_column(
"datasource_stages",
sa.Column(
"approval",
sa.JSON(),
server_default=sa.text("'{}'"),
nullable=False,
),
)
op.add_column(
"datasource_materializations",
sa.Column("disposed_at", sa.DateTime(timezone=True), nullable=True),
)
op.add_column(
"datasource_materializations",
sa.Column(
"disposition",
sa.JSON(),
server_default=sa.text("'{}'"),
nullable=False,
),
)
op.create_index(
"ix_datasource_materializations_disposed_at",
"datasource_materializations",
["disposed_at"],
unique=False,
)
op.create_table(
"datasource_lifecycle_evidence",
sa.Column("id", sa.String(length=36), nullable=False),
sa.Column("tenant_id", sa.String(length=36), nullable=False),
sa.Column("subject_ref", sa.String(length=160), nullable=False),
sa.Column("event_type", sa.String(length=80), nullable=False),
sa.Column("occurred_at", sa.DateTime(timezone=True), nullable=False),
sa.Column("actor_ref", sa.String(length=255), nullable=True),
sa.Column("policy_version", sa.String(length=120), nullable=True),
sa.Column("policy_hash", sa.String(length=64), nullable=True),
sa.Column("subject_digest", sa.String(length=64), nullable=False),
sa.Column("previous_event_hash", sa.String(length=64), nullable=True),
sa.Column("event_hash", sa.String(length=64), nullable=False),
sa.Column("details", sa.JSON(), nullable=False),
sa.PrimaryKeyConstraint("id"),
sa.UniqueConstraint(
"tenant_id",
"event_hash",
name="uq_datasource_lifecycle_evidence_hash",
),
)
op.create_index(
"ix_datasource_lifecycle_evidence_tenant_id",
"datasource_lifecycle_evidence",
["tenant_id"],
unique=False,
)
op.create_index(
"ix_datasource_lifecycle_evidence_subject_ref",
"datasource_lifecycle_evidence",
["subject_ref"],
unique=False,
)
op.create_index(
"ix_datasource_lifecycle_evidence_event_type",
"datasource_lifecycle_evidence",
["event_type"],
unique=False,
)
op.create_index(
"ix_datasource_lifecycle_evidence_actor_ref",
"datasource_lifecycle_evidence",
["actor_ref"],
unique=False,
)
op.create_index(
"ix_datasource_lifecycle_evidence_event_hash",
"datasource_lifecycle_evidence",
["event_hash"],
unique=False,
)
op.create_index(
"ix_datasource_lifecycle_evidence_subject",
"datasource_lifecycle_evidence",
["tenant_id", "subject_ref", "occurred_at"],
unique=False,
)
op.create_index(
"ix_datasource_lifecycle_evidence_event",
"datasource_lifecycle_evidence",
["tenant_id", "event_type", "occurred_at"],
unique=False,
)
def downgrade() -> None:
op.drop_table("datasource_lifecycle_evidence")
op.drop_index(
"ix_datasource_materializations_disposed_at",
table_name="datasource_materializations",
)
op.drop_column("datasource_materializations", "disposition")
op.drop_column("datasource_materializations", "disposed_at")
op.drop_column("datasource_stages", "approval")
op.drop_column("datasource_catalogue", "retention_policy")
op.drop_column("datasource_catalogue", "approval_policy")
@@ -9,6 +9,9 @@ from sqlalchemy import delete, func, insert, select
from sqlalchemy.orm import Session from sqlalchemy.orm import Session
from govoplan_core.core.datasources import ( from govoplan_core.core.datasources import (
DatasourceArtifactBackend,
DatasourceArtifactReference,
DatasourceField,
DatasourceUnavailableError, DatasourceUnavailableError,
DatasourceValidationError, DatasourceValidationError,
) )
@@ -116,6 +119,91 @@ class DatabaseRowsPayloadBackend:
) )
class ExternalArtifactPayloadBackend:
"""Adapts a Core artifact backend to persisted Datasources payload rows."""
def __init__(self, backend: DatasourceArtifactBackend) -> None:
self._backend = backend
self.backend = backend.backend
def read_rows(
self,
session: Session,
payload: DatasourcePayloadRecord,
*,
offset: int,
limit: int,
) -> Sequence[Mapping[str, object]]:
return self._backend.read_rows(
session,
tenant_id=payload.tenant_id,
artifact=_artifact_from_payload(payload),
offset=offset,
limit=limit,
)
def verify(
self,
session: Session,
payload: DatasourcePayloadRecord,
) -> None:
self._backend.verify(
session,
tenant_id=payload.tenant_id,
artifact=_artifact_from_payload(payload),
)
def delete(
self,
session: Session,
payload: DatasourcePayloadRecord,
) -> None:
self._backend.delete(
session,
tenant_id=payload.tenant_id,
artifact=_artifact_from_payload(payload),
)
def _artifact_from_payload(
payload: DatasourcePayloadRecord,
) -> DatasourceArtifactReference:
metadata = dict(payload.metadata_)
raw_schema = metadata.get("artifact_schema")
schema = tuple(
DatasourceField(
name=str(item.get("name") or ""),
data_type=str(item.get("data_type") or "unknown"),
nullable=bool(item.get("nullable", True)),
)
for item in raw_schema
if isinstance(item, Mapping)
) if isinstance(raw_schema, Sequence) and not isinstance(
raw_schema, (str, bytes)
) else ()
return DatasourceArtifactReference(
backend=payload.backend,
locator=str(payload.locator or ""),
checksum=payload.checksum,
row_count=payload.row_count,
byte_count=payload.byte_count,
schema=schema,
fingerprint=str(metadata.get("publication_fingerprint") or ""),
media_type=payload.media_type,
checkpoint=dict(payload.checkpoint_),
metadata={
str(key): value
for key, value in metadata.items()
if key not in {"artifact_schema", "artifact_validation"}
},
validation=(
dict(metadata.get("artifact_validation"))
if isinstance(metadata.get("artifact_validation"), Mapping)
else {}
),
)
class PayloadBackendRegistry: class PayloadBackendRegistry:
def __init__( def __init__(
self, self,
@@ -363,6 +451,7 @@ __all__ = [
"DATABASE_ROWS_BACKEND", "DATABASE_ROWS_BACKEND",
"DatabaseRowsPayloadBackend", "DatabaseRowsPayloadBackend",
"DatasourcePayloadBackend", "DatasourcePayloadBackend",
"ExternalArtifactPayloadBackend",
"PayloadBackendRegistry", "PayloadBackendRegistry",
"create_database_rows_payload", "create_database_rows_payload",
"create_external_payload_reference", "create_external_payload_reference",
+242 -1
View File
@@ -1,6 +1,9 @@
from __future__ import annotations from __future__ import annotations
import hashlib
import json
from collections.abc import Mapping from collections.abc import Mapping
from datetime import UTC, datetime
from fastapi import APIRouter, Depends, HTTPException, Query, status from fastapi import APIRouter, Depends, HTTPException, Query, status
from sqlalchemy.orm import Session from sqlalchemy.orm import Session
@@ -28,6 +31,8 @@ from govoplan_datasources.backend.schemas import (
DatasourceFreezeRequest, DatasourceFreezeRequest,
DatasourceGovernancePayload, DatasourceGovernancePayload,
DatasourceGovernanceUpdateRequest, DatasourceGovernanceUpdateRequest,
DatasourceLifecycleEvidenceListResponse,
DatasourceLifecycleEvidenceResponse,
DatasourceListResponse, DatasourceListResponse,
DatasourceMaterializationListResponse, DatasourceMaterializationListResponse,
DatasourceMaterializationResponse, DatasourceMaterializationResponse,
@@ -39,8 +44,13 @@ from govoplan_datasources.backend.schemas import (
DatasourcePreviewDiagnosticResponse, DatasourcePreviewDiagnosticResponse,
DatasourcePreviewResponse, DatasourcePreviewResponse,
DatasourceResponse, DatasourceResponse,
DatasourceRetentionApplyRequest,
DatasourceRetentionApplyResponse,
DatasourceRetentionCandidateResponse,
DatasourceRetentionPlanResponse,
DatasourceRetireResponse, DatasourceRetireResponse,
DatasourceStageCreateRequest, DatasourceStageCreateRequest,
DatasourceStageDecisionRequest,
DatasourceStageListResponse, DatasourceStageListResponse,
DatasourceStagePromoteRequest, DatasourceStagePromoteRequest,
DatasourceStagePromoteResponse, DatasourceStagePromoteResponse,
@@ -50,6 +60,7 @@ from govoplan_datasources.backend.service import (
ADMIN_SCOPE, ADMIN_SCOPE,
CATALOGUE_READ_SCOPE, CATALOGUE_READ_SCOPE,
SOURCE_WRITE_SCOPE, SOURCE_WRITE_SCOPE,
STAGE_APPROVE_SCOPE,
STAGE_WRITE_SCOPE, STAGE_WRITE_SCOPE,
SqlDatasourceProvider, SqlDatasourceProvider,
) )
@@ -234,6 +245,56 @@ def api_create_stage(
"schema_classification": _safe_mapping( "schema_classification": _safe_mapping(
stage.validation.get("schema_change") stage.validation.get("schema_change")
).get("classification"), ).get("classification"),
"approval_state": stage.approval.get("state"),
"approval_policy_hash": stage.approval.get("policy_hash"),
},
)
session.commit()
return _stage_response(stage)
@router.post(
"/stages/{stage_id}/decision",
response_model=DatasourceStageResponse,
)
def api_decide_stage(
stage_id: str,
payload: DatasourceStageDecisionRequest,
session: Session = Depends(get_session),
principal: ApiPrincipal = Depends(get_api_principal),
) -> DatasourceStageResponse:
_require_any_scope(principal, STAGE_APPROVE_SCOPE, ADMIN_SCOPE)
try:
stage, evidence_hash, replayed = _provider().decide_stage(
session,
principal,
stage_ref=f"stage:{stage_id}",
decision=payload.decision,
reason=payload.reason,
expected_policy_hash=payload.expected_policy_hash,
expected_subject_digest=payload.expected_subject_digest,
)
except DatasourceError as exc:
raise _http_error(exc) from exc
action = (
"datasources.stage.approved"
if payload.decision == "approve"
else "datasources.stage.rejected"
)
_audit(
session,
principal,
action=action,
object_type="datasource_stage",
object_id=stage.ref,
details={
"decision": payload.decision,
"resulting_state": stage.state,
"approval_count": stage.approval.get("approval_count"),
"required_approvals": stage.approval.get("required_approvals"),
"policy_hash": stage.approval.get("policy_hash"),
"evidence_hash": evidence_hash,
"replayed": replayed,
}, },
) )
session.commit() session.commit()
@@ -275,6 +336,12 @@ def api_promote_stage(
"quality_policy_hash": _safe_mapping( "quality_policy_hash": _safe_mapping(
materialization.provenance.get("stage_validation") materialization.provenance.get("stage_validation")
).get("policy_hash"), ).get("policy_hash"),
"approval_policy_hash": _safe_mapping(
materialization.provenance.get("stage_approval")
).get("policy_hash"),
"promotion_evidence_hash": materialization.provenance.get(
"promotion_evidence_hash"
),
}, },
) )
session.commit() session.commit()
@@ -284,6 +351,100 @@ def api_promote_stage(
) )
@router.get(
"/lifecycle-evidence",
response_model=DatasourceLifecycleEvidenceListResponse,
)
def api_list_lifecycle_evidence(
subject_ref: str | None = Query(default=None, max_length=160),
limit: int = Query(default=200, ge=1, le=500),
session: Session = Depends(get_session),
principal: ApiPrincipal = Depends(get_api_principal),
) -> DatasourceLifecycleEvidenceListResponse:
_require_any_scope(principal, CATALOGUE_READ_SCOPE, ADMIN_SCOPE)
try:
rows = _provider().list_lifecycle_evidence(
session,
principal,
subject_ref=subject_ref,
limit=limit,
)
except DatasourceError as exc:
raise _http_error(exc) from exc
return DatasourceLifecycleEvidenceListResponse(
evidence=[_evidence_response(item) for item in rows]
)
@router.get(
"/retention/plan",
response_model=DatasourceRetentionPlanResponse,
)
def api_preview_retention(
as_of: datetime | None = Query(default=None),
session: Session = Depends(get_session),
principal: ApiPrincipal = Depends(get_api_principal),
) -> DatasourceRetentionPlanResponse:
_require_any_scope(principal, ADMIN_SCOPE)
effective_at = as_of or datetime.now(UTC)
try:
plan = _provider().preview_retention(
session,
principal,
as_of=effective_at,
)
except DatasourceError as exc:
raise _http_error(exc) from exc
return DatasourceRetentionPlanResponse(
as_of=plan.as_of.isoformat(),
plan_hash=plan.plan_hash,
candidates=[
DatasourceRetentionCandidateResponse(**item.to_dict())
for item in plan.candidates
],
)
@router.post(
"/retention/apply",
response_model=DatasourceRetentionApplyResponse,
)
def api_apply_retention(
payload: DatasourceRetentionApplyRequest,
session: Session = Depends(get_session),
principal: ApiPrincipal = Depends(get_api_principal),
) -> DatasourceRetentionApplyResponse:
_require_any_scope(principal, ADMIN_SCOPE)
try:
disposed, evidence_hashes = _provider().apply_retention(
session,
principal,
as_of=payload.as_of,
plan_hash=payload.plan_hash,
target_refs=payload.target_refs,
)
except DatasourceError as exc:
raise _http_error(exc) from exc
_audit(
session,
principal,
action="datasources.retention.applied",
object_type="datasource_retention_plan",
object_id=payload.plan_hash,
details={
"as_of": payload.as_of.isoformat(),
"disposed_refs": list(disposed),
"evidence_hashes": list(evidence_hashes),
},
)
session.commit()
return DatasourceRetentionApplyResponse(
plan_hash=payload.plan_hash,
disposed_refs=list(disposed),
evidence_hashes=list(evidence_hashes),
)
@router.get("", response_model=DatasourceListResponse) @router.get("", response_model=DatasourceListResponse)
def api_list_datasources( def api_list_datasources(
query: str = Query(default="", max_length=200), query: str = Query(default="", max_length=200),
@@ -370,7 +531,17 @@ def api_update_datasource_governance(
action="datasources.governance.updated", action="datasources.governance.updated",
object_type="datasource", object_type="datasource",
object_id=item.ref, object_id=item.ref,
details={"governance": item.governance.to_dict()}, details={
"classification": item.governance.classification,
"publication_state": item.governance.publication_state,
"access_policy_ref": item.governance.access_policy_ref,
"visibility_policy_configured": bool(
item.governance.visibility_policy
),
"visibility_policy_hash": _mapping_hash(
item.governance.visibility_policy
),
},
) )
session.commit() session.commit()
return _datasource_response(item) return _datasource_response(item)
@@ -495,6 +666,46 @@ def api_refresh_datasource(
) )
@router.post(
"/{datasource_id}/refresh/stage",
response_model=DatasourceStageResponse,
status_code=status.HTTP_201_CREATED,
)
def api_prepare_refresh(
datasource_id: str,
session: Session = Depends(get_session),
principal: ApiPrincipal = Depends(get_api_principal),
) -> DatasourceStageResponse:
_require_any_scope(
principal,
SOURCE_WRITE_SCOPE,
STAGE_WRITE_SCOPE,
ADMIN_SCOPE,
)
try:
stage = _provider().prepare_refresh(
session,
principal,
datasource_ref=f"datasource:{datasource_id}",
)
except DatasourceError as exc:
raise _http_error(exc) from exc
_audit(
session,
principal,
action="datasources.refresh.staged",
object_type="datasource_stage",
object_id=stage.ref,
details={
"datasource_ref": stage.target_datasource_ref,
"approval_state": stage.approval.get("state"),
"quality_policy_hash": stage.validation.get("policy_hash"),
},
)
session.commit()
return _stage_response(stage)
@router.post( @router.post(
"/{datasource_id}/freeze", "/{datasource_id}/freeze",
response_model=DatasourceMaterializationResponse, response_model=DatasourceMaterializationResponse,
@@ -580,6 +791,7 @@ def _datasource_response(item: DatasourceDescriptor) -> DatasourceResponse:
name=field.name, name=field.name,
data_type=field.data_type, data_type=field.data_type,
nullable=field.nullable, nullable=field.nullable,
classification=field.classification,
) )
for field in item.schema for field in item.schema
], ],
@@ -610,6 +822,7 @@ def _materialization_response(
name=field.name, name=field.name,
data_type=field.data_type, data_type=field.data_type,
nullable=field.nullable, nullable=field.nullable,
classification=field.classification,
) )
for field in item.schema for field in item.schema
], ],
@@ -621,6 +834,8 @@ def _materialization_response(
item.source_timestamp.isoformat() if item.source_timestamp else None item.source_timestamp.isoformat() if item.source_timestamp else None
), ),
created_at=item.created_at.isoformat() if item.created_at else None, created_at=item.created_at.isoformat() if item.created_at else None,
disposed_at=item.disposed_at.isoformat() if item.disposed_at else None,
disposition=dict(item.disposition),
provenance=dict(item.provenance), provenance=dict(item.provenance),
metadata=dict(item.metadata), metadata=dict(item.metadata),
governance=item.governance.to_dict(), governance=item.governance.to_dict(),
@@ -643,12 +858,14 @@ def _stage_response(item: DatasourceStage) -> DatasourceStageResponse:
name=field.name, name=field.name,
data_type=field.data_type, data_type=field.data_type,
nullable=field.nullable, nullable=field.nullable,
classification=field.classification,
) )
for field in item.schema for field in item.schema
], ],
row_count=item.row_count, row_count=item.row_count,
byte_count=item.byte_count, byte_count=item.byte_count,
validation=dict(item.validation), validation=dict(item.validation),
approval=dict(item.approval),
created_at=item.created_at.isoformat() if item.created_at else None, created_at=item.created_at.isoformat() if item.created_at else None,
promoted_at=item.promoted_at.isoformat() if item.promoted_at else None, promoted_at=item.promoted_at.isoformat() if item.promoted_at else None,
promoted_materialization_ref=item.promoted_materialization_ref, promoted_materialization_ref=item.promoted_materialization_ref,
@@ -673,6 +890,7 @@ def _origin_response(item: DatasourceOrigin) -> DatasourceOriginResponse:
name=field.name, name=field.name,
data_type=field.data_type, data_type=field.data_type,
nullable=field.nullable, nullable=field.nullable,
classification=field.classification,
) )
for field in item.schema for field in item.schema
], ],
@@ -703,6 +921,22 @@ def _origin_response(item: DatasourceOrigin) -> DatasourceOriginResponse:
) )
def _evidence_response(item) -> DatasourceLifecycleEvidenceResponse:
return DatasourceLifecycleEvidenceResponse(
ref=f"lifecycle-evidence:{item.id}",
subject_ref=item.subject_ref,
event_type=item.event_type,
occurred_at=item.occurred_at.isoformat(),
actor_ref=item.actor_ref,
policy_version=item.policy_version,
policy_hash=item.policy_hash,
subject_digest=item.subject_digest,
previous_event_hash=item.previous_event_hash,
event_hash=item.event_hash,
details=dict(item.details_),
)
def _governance( def _governance(
payload: DatasourceGovernancePayload | None, payload: DatasourceGovernancePayload | None,
) -> DatasourceGovernance | None: ) -> DatasourceGovernance | None:
@@ -715,6 +949,13 @@ def _safe_mapping(value: object) -> Mapping[str, object]:
return value if isinstance(value, Mapping) else {} return value if isinstance(value, Mapping) else {}
def _mapping_hash(value: Mapping[str, object]) -> str | None:
if not value:
return None
encoded = json.dumps(value, sort_keys=True, separators=(",", ":"), default=str)
return hashlib.sha256(encoded.encode("utf-8")).hexdigest()
def _audit( def _audit(
session: Session, session: Session,
principal: ApiPrincipal, principal: ApiPrincipal,
@@ -1,5 +1,6 @@
from __future__ import annotations from __future__ import annotations
from datetime import datetime
from typing import Any, Literal from typing import Any, Literal
from pydantic import BaseModel, Field, model_validator from pydantic import BaseModel, Field, model_validator
@@ -43,11 +44,15 @@ class DatasourceGovernancePayload(BaseModel):
classification: str = Field(default="internal", min_length=1, max_length=80) classification: str = Field(default="internal", min_length=1, max_length=80)
privacy_profile_ref: str | None = Field(default=None, max_length=500) privacy_profile_ref: str | None = Field(default=None, max_length=500)
retention_policy_ref: str | None = Field(default=None, max_length=500) retention_policy_ref: str | None = Field(default=None, max_length=500)
access_policy_ref: str | None = Field(default=None, max_length=500)
visibility_policy: dict[str, Any] = Field(default_factory=dict)
hold_refs: list[str] = Field(default_factory=list, max_length=100) hold_refs: list[str] = Field(default_factory=list, max_length=100)
publication_state: str = Field(default="draft", min_length=1, max_length=50) publication_state: str = Field(default="draft", min_length=1, max_length=50)
transfer_agreement_ref: str | None = Field(default=None, max_length=500) transfer_agreement_ref: str | None = Field(default=None, max_length=500)
freshness_policy: dict[str, Any] = Field(default_factory=dict) freshness_policy: dict[str, Any] = Field(default_factory=dict)
quality_policy: dict[str, Any] = Field(default_factory=dict) quality_policy: dict[str, Any] = Field(default_factory=dict)
approval_policy: dict[str, Any] = Field(default_factory=dict)
retention_policy: dict[str, Any] = Field(default_factory=dict)
known_limits: list[str] = Field(default_factory=list, max_length=100) known_limits: list[str] = Field(default_factory=list, max_length=100)
correction_procedure_ref: str | None = Field(default=None, max_length=500) correction_procedure_ref: str | None = Field(default=None, max_length=500)
affected_refs: list[str] = Field(default_factory=list, max_length=250) affected_refs: list[str] = Field(default_factory=list, max_length=250)
@@ -70,6 +75,7 @@ class DatasourceFieldResponse(BaseModel):
name: str name: str
data_type: str data_type: str
nullable: bool nullable: bool
classification: str = "internal"
class DatasourceResponse(BaseModel): class DatasourceResponse(BaseModel):
@@ -113,6 +119,8 @@ class DatasourceMaterializationResponse(BaseModel):
frozen_label: str | None frozen_label: str | None
source_timestamp: str | None source_timestamp: str | None
created_at: str | None created_at: str | None
disposed_at: str | None
disposition: dict[str, Any]
provenance: dict[str, Any] provenance: dict[str, Any]
metadata: dict[str, Any] metadata: dict[str, Any]
governance: DatasourceGovernancePayload governance: DatasourceGovernancePayload
@@ -178,6 +186,7 @@ class DatasourceStageResponse(BaseModel):
row_count: int | None row_count: int | None
byte_count: int | None byte_count: int | None
validation: DatasourceStageValidationResponse validation: DatasourceStageValidationResponse
approval: dict[str, Any]
created_at: str | None created_at: str | None
promoted_at: str | None promoted_at: str | None
promoted_materialization_ref: str | None promoted_materialization_ref: str | None
@@ -220,6 +229,61 @@ class DatasourceStagePromoteRequest(BaseModel):
frozen_label: str | None = Field(default=None, max_length=300) frozen_label: str | None = Field(default=None, max_length=300)
class DatasourceStageDecisionRequest(BaseModel):
decision: Literal["approve", "reject"]
reason: str = Field(min_length=1, max_length=2_000)
expected_policy_hash: str = Field(min_length=64, max_length=64)
expected_subject_digest: str = Field(min_length=64, max_length=64)
class DatasourceLifecycleEvidenceResponse(BaseModel):
ref: str
subject_ref: str
event_type: str
occurred_at: str
actor_ref: str | None
policy_version: str | None
policy_hash: str | None
subject_digest: str
previous_event_hash: str | None
event_hash: str
details: dict[str, Any]
class DatasourceLifecycleEvidenceListResponse(BaseModel):
evidence: list[DatasourceLifecycleEvidenceResponse]
class DatasourceRetentionCandidateResponse(BaseModel):
ref: str
kind: Literal["stage", "materialization"]
datasource_ref: str | None
disposition: Literal["delete_stage", "purge_materialization_payload"]
eligible_at: str
eligible: bool
blockers: list[str]
policy_version: str
policy_hash: str
class DatasourceRetentionPlanResponse(BaseModel):
as_of: str
plan_hash: str
candidates: list[DatasourceRetentionCandidateResponse]
class DatasourceRetentionApplyRequest(BaseModel):
as_of: datetime
plan_hash: str = Field(min_length=64, max_length=64)
target_refs: list[str] = Field(min_length=1, max_length=500)
class DatasourceRetentionApplyResponse(BaseModel):
plan_hash: str
disposed_refs: list[str]
evidence_hashes: list[str]
class DatasourceStagePromoteResponse(BaseModel): class DatasourceStagePromoteResponse(BaseModel):
datasource: DatasourceResponse datasource: DatasourceResponse
materialization: DatasourceMaterializationResponse materialization: DatasourceMaterializationResponse
@@ -0,0 +1,179 @@
from __future__ import annotations
from collections.abc import Mapping, Sequence
from urllib.parse import quote
from sqlalchemy import func, select
from sqlalchemy.orm import Session
from govoplan_core.auth import ApiPrincipal
from govoplan_core.core.modules import ModuleContext
from govoplan_core.core.search import (
SearchAuthorizationRequest,
SearchBackfillPage,
SearchBackfillRequest,
SearchDocument,
SearchResourceType,
)
from govoplan_datasources.backend.db.models import DatasourceRecord
from govoplan_datasources.backend.service import ADMIN_SCOPE, CATALOGUE_READ_SCOPE
PROVIDER_ID = "datasources.catalogue"
RESOURCE_TYPE = "datasource"
class DatasourcesSearchSource:
def resource_types(self) -> Sequence[SearchResourceType]:
return (
SearchResourceType(
provider_id=PROVIDER_ID,
module_id="datasources",
resource_type=RESOURCE_TYPE,
label="Datasources",
requires_authorization_recheck=True,
),
)
def backfill(
self,
session: object,
*,
request: SearchBackfillRequest,
) -> SearchBackfillPage:
_assert_source(request.provider_id, request.resource_type)
db = _session(session)
statement = select(DatasourceRecord).where(
DatasourceRecord.tenant_id == request.tenant_id,
DatasourceRecord.deleted_at.is_(None),
)
if request.cursor:
statement = statement.where(DatasourceRecord.id > request.cursor)
rows = list(
db.scalars(
statement.order_by(DatasourceRecord.id).limit(request.limit + 1)
).all()
)
has_more = len(rows) > request.limit
selected = rows[: request.limit]
high_watermark = db.scalar(
select(func.max(DatasourceRecord.updated_at)).where(
DatasourceRecord.tenant_id == request.tenant_id,
DatasourceRecord.deleted_at.is_(None),
)
)
return SearchBackfillPage(
documents=tuple(_document(row) for row in selected),
next_cursor=selected[-1].id if has_more and selected else None,
complete=not has_more,
high_watermark=high_watermark.isoformat() if high_watermark else None,
)
def authorize(
self,
session: object,
principal: object,
*,
requests: Sequence[SearchAuthorizationRequest],
) -> Mapping[str, bool]:
decisions = {item.reference.key: False for item in requests}
if not isinstance(principal, ApiPrincipal) or not (
principal.has(CATALOGUE_READ_SCOPE) or principal.has(ADMIN_SCOPE)
):
return decisions
db = _session(session)
eligible = [
request
for request in requests
if request.reference.tenant_id == principal.tenant_id
and request.reference.module_id == "datasources"
and request.reference.resource_type == RESOURCE_TYPE
]
resource_ids = {request.reference.resource_id for request in eligible}
available_ids = (
set(
db.scalars(
select(DatasourceRecord.id).where(
DatasourceRecord.tenant_id == principal.tenant_id,
DatasourceRecord.id.in_(resource_ids),
DatasourceRecord.deleted_at.is_(None),
)
).all()
)
if resource_ids
else set()
)
for request in eligible:
decisions[request.reference.key] = (
request.reference.resource_id in available_ids
)
return decisions
def create_datasources_search_source(
_context: ModuleContext,
) -> DatasourcesSearchSource:
return DatasourcesSearchSource()
def _document(row: DatasourceRecord) -> SearchDocument:
datasource_ref = quote(f"datasource:{row.id}", safe="")
return SearchDocument(
tenant_id=row.tenant_id,
module_id="datasources",
provider_id=PROVIDER_ID,
resource_type=RESOURCE_TYPE,
resource_id=row.id,
title=row.name,
url=f"/datasources?datasource={datasource_ref}",
summary=(row.description or row.source_name)[:4000],
keywords=tuple(
value[:200]
for value in (
row.source_name,
row.kind,
row.mode,
row.shape,
row.status,
row.classification,
row.publication_state,
)
if value
),
visibility="restricted",
acl_tokens=(
f"scope:{CATALOGUE_READ_SCOPE}",
f"scope:{ADMIN_SCOPE}",
),
metadata={
"kind": row.kind,
"mode": row.mode,
"shape": row.shape,
"status": row.status,
"classification": row.classification,
"publication_state": row.publication_state,
"authority_mode": row.authority_mode,
},
source_revision=f"{row.schema_version}:{row.updated_at.isoformat()}",
source_updated_at=row.updated_at,
requires_authorization_recheck=True,
)
def _assert_source(provider_id: str, resource_type: str) -> None:
if provider_id != PROVIDER_ID or resource_type != RESOURCE_TYPE:
raise ValueError("Unsupported Datasources search source.")
def _session(value: object) -> Session:
if not isinstance(value, Session):
raise TypeError("Datasources search requires a SQLAlchemy session.")
return value
__all__ = [
"DatasourcesSearchSource",
"PROVIDER_ID",
"RESOURCE_TYPE",
"create_datasources_search_source",
]
File diff suppressed because it is too large Load Diff
@@ -145,6 +145,7 @@ def field_payload(field: DatasourceField) -> dict[str, object]:
"name": field.name, "name": field.name,
"data_type": field.data_type, "data_type": field.data_type,
"nullable": field.nullable, "nullable": field.nullable,
"classification": field.classification,
} }
@@ -0,0 +1,490 @@
from __future__ import annotations
import hashlib
import json
from collections.abc import Mapping, Sequence
from dataclasses import dataclass
from typing import cast
from govoplan_core.core.access import PrincipalRef
from govoplan_core.core.datasources import (
DatasourceAccessError,
DatasourceField,
DatasourceValidationError,
)
from govoplan_core.core.tabular_sources import TabularPreviewDiagnostic
_ACL_KEYS = frozenset(
{
"account_ids",
"membership_ids",
"identity_ids",
"group_ids",
"role_ids",
"service_account_ids",
"auth_methods",
}
)
_CLAIM_KEYS = frozenset(
{
"account_id",
"membership_id",
"identity_id",
"service_account_id",
"group_ids",
"role_ids",
"function_assignment_ids",
}
)
_POLICY_KEYS = frozenset({"source_acl", "materialization_acl", "fields", "row_filters"})
@dataclass(frozen=True, slots=True)
class FieldRule:
name: str
classification: str
action: str
allowed: bool
@dataclass(frozen=True, slots=True)
class RowFilter:
field: str
claim: str
operator: str
allow_null: bool
@dataclass(frozen=True, slots=True)
class VisibilityPlan:
applied: bool
policy_hashes: tuple[str, ...]
field_rules: tuple[FieldRule, ...]
row_filters: tuple[RowFilter, ...]
principal_fingerprint: str
decision_refs: tuple[str, ...] = ()
def view_fingerprint(self, base_fingerprint: str) -> str:
if not self.applied:
return base_fingerprint
return _digest(
{
"base_fingerprint": base_fingerprint,
"policy_hashes": self.policy_hashes,
"principal_fingerprint": self.principal_fingerprint,
"decision_refs": self.decision_refs,
}
)
def visibility_plan(
*,
principal: PrincipalRef,
policies: Sequence[Mapping[str, object]],
materialized: bool,
schema: Sequence[DatasourceField],
decision_refs: Sequence[str] = (),
) -> VisibilityPlan:
normalized = tuple(_normalize_policy(policy) for policy in policies if policy)
if not normalized:
return VisibilityPlan(
applied=False,
policy_hashes=(),
field_rules=(),
row_filters=(),
principal_fingerprint="",
)
schema_names = {field.name for field in schema}
field_rules: list[FieldRule] = []
row_filters: list[RowFilter] = []
for policy in normalized:
source_acl = policy.get("source_acl")
if isinstance(source_acl, Mapping) and not _acl_allows(source_acl, principal):
raise DatasourceAccessError("Datasource visibility policy denied access.")
materialization_acl = policy.get("materialization_acl")
if (
materialized
and isinstance(materialization_acl, Mapping)
and not _acl_allows(materialization_acl, principal)
):
raise DatasourceAccessError(
"Datasource materialization policy denied access."
)
fields = policy.get("fields", {})
if isinstance(fields, Mapping):
for name, raw_rule in fields.items():
if name not in schema_names:
raise DatasourceValidationError(
f"Visibility policy references unknown field {name!r}."
)
assert isinstance(raw_rule, Mapping)
field_rules.append(
FieldRule(
name=name,
classification=str(raw_rule["classification"]),
action=str(raw_rule["action"]),
allowed=_acl_allows(raw_rule["allow"], principal),
)
)
for raw_filter in policy.get("row_filters", ()):
assert isinstance(raw_filter, Mapping)
field = str(raw_filter["field"])
if field not in schema_names:
raise DatasourceValidationError(
f"Visibility policy references unknown row-filter field {field!r}."
)
row_filters.append(
RowFilter(
field=field,
claim=str(raw_filter["claim"]),
operator=str(raw_filter["operator"]),
allow_null=bool(raw_filter["allow_null"]),
)
)
all_facts = _principal_policy_facts(principal)
relevant_fact_keys = _relevant_policy_fact_keys(normalized)
principal_payload = {key: all_facts[key] for key in sorted(relevant_fact_keys)}
return VisibilityPlan(
applied=True,
policy_hashes=tuple(_digest(policy) for policy in normalized),
field_rules=tuple(field_rules),
row_filters=tuple(row_filters),
principal_fingerprint=_digest(principal_payload),
decision_refs=tuple(sorted(set(decision_refs))),
)
def normalize_visibility_policy(
policy: Mapping[str, object],
*,
schema: Sequence[DatasourceField] = (),
) -> Mapping[str, object]:
normalized = _normalize_policy(policy)
if schema:
known = {field.name for field in schema}
configured = set(cast(Mapping[str, object], normalized.get("fields", {})))
configured.update(
str(item["field"])
for item in cast(
Sequence[Mapping[str, object]], normalized.get("row_filters", ())
)
)
unknown = sorted(configured - known)
if unknown:
raise DatasourceValidationError(
f"Visibility policy references unknown fields: {', '.join(unknown)}."
)
return normalized
def visible_schema(
schema: Sequence[DatasourceField],
plan: VisibilityPlan,
) -> tuple[DatasourceField, ...]:
classifications: dict[str, str] = {}
omitted: set[str] = set()
for rule in plan.field_rules:
classifications[rule.name] = rule.classification
if not rule.allowed and rule.action == "omit":
omitted.add(rule.name)
return tuple(
DatasourceField(
name=field.name,
data_type=field.data_type,
nullable=field.nullable,
classification=classifications.get(field.name, field.classification),
)
for field in schema
if field.name not in omitted
)
def apply_visibility(
rows: Sequence[Mapping[str, object]],
*,
principal: PrincipalRef,
plan: VisibilityPlan,
columns: Sequence[str] = (),
) -> tuple[tuple[Mapping[str, object], ...], int]:
omitted = {
rule.name
for rule in plan.field_rules
if not rule.allowed and rule.action == "omit"
}
redacted = {
rule.name
for rule in plan.field_rules
if not rule.allowed and rule.action == "redact"
}
requested = tuple(dict.fromkeys(columns))
result: list[Mapping[str, object]] = []
filtered = 0
for source_row in rows:
if not all(
_row_filter_allows(source_row, item, principal) for item in plan.row_filters
):
filtered += 1
continue
names = requested or tuple(source_row)
result.append(
{
name: None if name in redacted else source_row.get(name)
for name in names
if name not in omitted
}
)
return tuple(result), filtered
def visibility_diagnostics(
plan: VisibilityPlan,
*,
scanned_rows: int,
permitted_rows: int,
filtered_rows: int,
scan_limited: bool,
) -> tuple[TabularPreviewDiagnostic, ...]:
if not plan.applied:
return ()
diagnostics = [
TabularPreviewDiagnostic(
severity="info",
code="datasource.visibility_applied",
message="Datasource visibility policy was applied before rows left the provider.",
details={
"policy_count": len(plan.policy_hashes),
"field_rule_count": len(plan.field_rules),
"row_filter_count": len(plan.row_filters),
"scanned_rows": scanned_rows,
"permitted_rows": permitted_rows,
"filtered_rows": filtered_rows,
},
)
]
if scan_limited:
diagnostics.append(
TabularPreviewDiagnostic(
severity="warning",
code="datasource.visibility_scan_limited",
message="The bounded visibility scan ended before the origin was exhausted.",
details={"scanned_rows": scanned_rows},
)
)
return tuple(diagnostics)
def _normalize_policy(policy: Mapping[str, object]) -> dict[str, object]:
unknown = set(policy) - _POLICY_KEYS
if unknown:
raise DatasourceValidationError(
f"Unsupported visibility policy keys: {', '.join(sorted(unknown))}."
)
result: dict[str, object] = {}
for name in ("source_acl", "materialization_acl"):
if name in policy:
result[name] = _normalize_acl(policy[name], path=name)
raw_fields = policy.get("fields", {})
if not isinstance(raw_fields, Mapping) or len(raw_fields) > 500:
raise DatasourceValidationError(
"Visibility policy fields must be a mapping of at most 500 fields."
)
fields: dict[str, object] = {}
for raw_name, raw_rule in sorted(raw_fields.items(), key=lambda item: str(item[0])):
name = str(raw_name).strip()
if not name or not isinstance(raw_rule, Mapping):
raise DatasourceValidationError(
"Every visibility field rule needs a field name and mapping."
)
unknown_rule = set(raw_rule) - {"classification", "action", "allow"}
if unknown_rule:
raise DatasourceValidationError(
f"Unsupported visibility field keys for {name!r}: {', '.join(sorted(unknown_rule))}."
)
classification = str(raw_rule.get("classification") or "restricted").strip()
if not classification or len(classification) > 80:
raise DatasourceValidationError(
"Field classifications must contain 1 to 80 characters."
)
action = str(raw_rule.get("action") or "redact").strip()
if action not in {"redact", "omit"}:
raise DatasourceValidationError(
"Field visibility action must be redact or omit."
)
if "allow" not in raw_rule:
raise DatasourceValidationError(
f"Visibility field rule {name!r} needs an allow ACL."
)
fields[name] = {
"classification": classification,
"action": action,
"allow": _normalize_acl(raw_rule["allow"], path=f"fields.{name}.allow"),
}
if fields:
result["fields"] = fields
raw_filters = policy.get("row_filters", ())
if not isinstance(raw_filters, Sequence) or isinstance(raw_filters, (str, bytes)):
raise DatasourceValidationError("Visibility row_filters must be a list.")
if len(raw_filters) > 50:
raise DatasourceValidationError(
"Visibility policy supports at most 50 row filters."
)
filters: list[dict[str, object]] = []
for raw_filter in raw_filters:
if not isinstance(raw_filter, Mapping):
raise DatasourceValidationError(
"Every visibility row filter must be a mapping."
)
unknown_filter = set(raw_filter) - {"field", "claim", "operator", "allow_null"}
if unknown_filter:
raise DatasourceValidationError(
f"Unsupported row-filter keys: {', '.join(sorted(unknown_filter))}."
)
field = str(raw_filter.get("field") or "").strip()
claim = str(raw_filter.get("claim") or "").strip()
operator = str(raw_filter.get("operator") or "equals").strip()
if not field:
raise DatasourceValidationError(
"Every visibility row filter needs a field."
)
if claim not in _CLAIM_KEYS:
raise DatasourceValidationError(
f"Unsupported visibility row-filter claim {claim!r}."
)
if operator not in {"equals", "in"}:
raise DatasourceValidationError("Row-filter operator must be equals or in.")
filters.append(
{
"field": field,
"claim": claim,
"operator": operator,
"allow_null": bool(raw_filter.get("allow_null", False)),
}
)
if filters:
result["row_filters"] = filters
return result
def _normalize_acl(value: object, *, path: str) -> dict[str, list[str]]:
if not isinstance(value, Mapping):
raise DatasourceValidationError(
f"Visibility policy {path} must be an ACL mapping."
)
unknown = set(value) - _ACL_KEYS
if unknown:
raise DatasourceValidationError(
f"Unsupported ACL selectors in {path}: {', '.join(sorted(unknown))}."
)
result: dict[str, list[str]] = {}
for key in sorted(value):
raw_values = value[key]
if not isinstance(raw_values, Sequence) or isinstance(raw_values, (str, bytes)):
raise DatasourceValidationError(
f"Visibility ACL selector {path}.{key} must be a list."
)
if len(raw_values) > 250:
raise DatasourceValidationError(
f"Visibility ACL selector {path}.{key} is limited to 250 entries."
)
normalized = sorted(
{str(item).strip() for item in raw_values if str(item).strip()}
)
result[key] = normalized
return result
def _acl_allows(acl: Mapping[str, object], principal: PrincipalRef) -> bool:
facts = _principal_policy_facts(principal)
return any(set(values) & set(facts.get(key, ())) for key, values in acl.items())
def _principal_policy_facts(principal: PrincipalRef) -> dict[str, tuple[str, ...]]:
return {
"account_ids": _optional_tuple(principal.account_id),
"membership_ids": _optional_tuple(principal.membership_id),
"identity_ids": _optional_tuple(principal.identity_id),
"group_ids": tuple(sorted(principal.group_ids)),
"role_ids": tuple(sorted(principal.role_ids)),
"service_account_ids": _optional_tuple(principal.service_account_id),
"auth_methods": (principal.auth_method,),
"function_assignment_ids": tuple(sorted(principal.function_assignment_ids)),
}
def _claim_values(principal: PrincipalRef, claim: str) -> tuple[str, ...]:
facts = _principal_policy_facts(principal)
aliases = {
"account_id": "account_ids",
"membership_id": "membership_ids",
"identity_id": "identity_ids",
"service_account_id": "service_account_ids",
}
return facts.get(aliases.get(claim, claim), ())
def _relevant_policy_fact_keys(
policies: Sequence[Mapping[str, object]],
) -> set[str]:
keys: set[str] = set()
aliases = {
"account_id": "account_ids",
"membership_id": "membership_ids",
"identity_id": "identity_ids",
"service_account_id": "service_account_ids",
}
for policy in policies:
for acl_name in ("source_acl", "materialization_acl"):
acl = policy.get(acl_name)
if isinstance(acl, Mapping):
keys.update(str(key) for key in acl)
fields = policy.get("fields")
if isinstance(fields, Mapping):
for rule in fields.values():
if isinstance(rule, Mapping) and isinstance(rule.get("allow"), Mapping):
keys.update(str(key) for key in rule["allow"])
for row_filter in policy.get("row_filters", ()):
if isinstance(row_filter, Mapping):
claim = str(row_filter.get("claim") or "")
keys.add(aliases.get(claim, claim))
return keys
def _row_filter_allows(
row: Mapping[str, object],
item: RowFilter,
principal: PrincipalRef,
) -> bool:
value = row.get(item.field)
if value is None:
return item.allow_null
claims = set(_claim_values(principal, item.claim))
if not claims:
return False
if isinstance(value, Sequence) and not isinstance(value, (str, bytes)):
values = {str(candidate) for candidate in value}
return bool(values & claims)
return str(value) in claims
def _optional_tuple(value: object | None) -> tuple[str, ...]:
return (str(value),) if value is not None and str(value) else ()
def _digest(value: object) -> str:
encoded = json.dumps(value, sort_keys=True, separators=(",", ":"), default=str)
return hashlib.sha256(encoded.encode("utf-8")).hexdigest()
__all__ = [
"VisibilityPlan",
"apply_visibility",
"normalize_visibility_policy",
"visibility_diagnostics",
"visibility_plan",
"visible_schema",
]
+630
View File
@@ -0,0 +1,630 @@
from __future__ import annotations
import json
import unittest
from datetime import UTC, datetime
from sqlalchemy import create_engine
from sqlalchemy.orm import Session
from govoplan_core.core.dsar import (
DsarErasureActionRef,
DsarProvider,
DsarRecordRef,
DsarSubjectRef,
)
from govoplan_core.db.base import Base
from govoplan_core.privacy.dsar_workflow import (
create_data_subject_request,
search_data_subject_request,
)
from govoplan_datasources.backend.db.models import (
DatasourceGovernanceReferenceRecord,
DatasourceLifecycleEvidenceRecord,
DatasourceMaterializationRecord,
DatasourcePayloadRecord,
DatasourcePayloadRowRecord,
DatasourcePublicationRecord,
DatasourceRecord,
DatasourceStageRecord,
)
from govoplan_datasources.backend.dsar_provider import (
DATASOURCES_DSAR_CAPABILITY,
DatasourcesDsarProvider,
)
from govoplan_datasources.backend.manifest import manifest
NOW = datetime(2026, 8, 21, 19, 0, tzinfo=UTC)
SECRET = "private-datasource-detail-do-not-export"
class _Registry:
def __init__(
self,
provider: DatasourcesDsarProvider,
*,
active: bool = True,
) -> None:
self.provider = provider
self.active = active
def capability_names(self):
return (DATASOURCES_DSAR_CAPABILITY,)
def capability_owner(self, name):
self._assert_capability(name)
return "datasources"
def tenant_entitlement_resolver(self):
active = self.active
class _Resolver:
@staticmethod
def resolve(session, tenant_id):
del session, tenant_id
return type(
"State",
(),
{"effective_modules": ("datasources",) if active else ()},
)()
return _Resolver()
def require_tenant_capability(self, name, session, **kwargs):
del session, kwargs
self._assert_capability(name)
return self.provider
def manifests(self):
return (type("Manifest", (), {"id": "datasources"})(),)
@staticmethod
def _assert_capability(name: str) -> None:
if name != DATASOURCES_DSAR_CAPABILITY:
raise KeyError(name)
class DatasourcesDsarProviderTests(unittest.TestCase):
def setUp(self) -> None:
self.engine = create_engine("sqlite+pysqlite:///:memory:")
Base.metadata.create_all(self.engine)
self.session = Session(self.engine)
self.provider = DatasourcesDsarProvider()
self.assertIsInstance(self.provider, DsarProvider)
self._seed()
self.session.commit()
def tearDown(self) -> None:
self.session.close()
self.engine.dispose()
def _seed(self) -> None:
datasource = self._datasource("datasource-1", "tenant-1")
other = self._datasource("datasource-other", "tenant-2")
self.session.add_all((datasource, other))
self.session.flush()
referenced_payload = self._payload(
"payload-referenced",
"tenant-1",
created_by="account-1",
)
orphan_payload = self._payload(
"payload-orphan",
"tenant-1",
created_by="account-1",
)
self.session.add_all((referenced_payload, orphan_payload))
self.session.flush()
self.session.add(
DatasourcePayloadRowRecord(
payload_id=referenced_payload.id,
row_index=0,
row_={"secret": SECRET},
checksum="a" * 64,
)
)
materialization = DatasourceMaterializationRecord(
id="materialization-1",
tenant_id="tenant-1",
datasource_id=datasource.id,
revision=1,
state="published",
schema_version=1,
schema_=[{"secret": SECRET}],
payload_id=referenced_payload.id,
payload_checksum="b" * 64,
rows=[{"secret": SECRET}],
fingerprint="c" * 64,
row_count=1,
byte_count=100,
source_timestamp=NOW,
provenance_={"secret": SECRET},
metadata_={"secret": SECRET},
governance_snapshot_={"secret": SECRET},
created_by="account-1",
)
self.session.add(materialization)
self.session.flush()
datasource.current_materialization_id = materialization.id
self.session.add_all(
(
DatasourceGovernanceReferenceRecord(
id="governance-1",
tenant_id="tenant-1",
datasource_id=datasource.id,
relation="authoritative_source",
reference=SECRET,
),
self._stage(
"stage-1",
target_datasource_id=datasource.id,
promoted=False,
),
self._stage(
"stage-promoted",
target_datasource_id=datasource.id,
promoted=True,
),
DatasourcePublicationRecord(
id="publication-1",
tenant_id="tenant-1",
producer_module="dataflow",
producer_run_ref=SECRET,
idempotency_key=SECRET,
request_hash="d" * 64,
datasource_id=datasource.id,
materialization_id=materialization.id,
status="published",
details_={"secret": SECRET},
created_by="account-1",
),
)
)
@staticmethod
def _datasource(row_id: str, tenant_id: str) -> DatasourceRecord:
return DatasourceRecord(
id=row_id,
tenant_id=tenant_id,
source_name=f"source-{row_id}",
name=SECRET,
description=SECRET,
kind="table",
mode="cached",
shape="tabular",
status="active",
provider="connector.provider",
provider_ref=SECRET,
schema_version=1,
schema_=[{"secret": SECRET}],
fingerprint="e" * 64,
row_count=1,
byte_count=100,
provenance_={"secret": SECRET},
metadata_={"secret": SECRET},
owner_ref=SECRET,
steward_ref=SECRET,
authoritative_source_ref=SECRET,
authority_mode="external_mirror",
legal_basis_refs=[SECRET],
purposes=[SECRET],
semantic_definition=SECRET,
official_keys=[SECRET],
classification="personal",
privacy_profile_ref=SECRET,
retention_policy_ref=SECRET,
hold_refs=[SECRET],
publication_state="published",
transfer_agreement_ref=SECRET,
freshness_policy={"secret": SECRET},
quality_policy={"secret": SECRET},
known_limits=[SECRET],
correction_procedure_ref=SECRET,
affected_refs=[SECRET],
dependency_refs=[SECRET],
created_by="account-1",
updated_by="account-1",
)
@staticmethod
def _payload(
row_id: str,
tenant_id: str,
*,
created_by: str,
) -> DatasourcePayloadRecord:
return DatasourcePayloadRecord(
id=row_id,
tenant_id=tenant_id,
backend="database_rows",
state="published",
locator=SECRET,
media_type="application/x-ndjson",
checksum="f" * 64,
row_count=1,
byte_count=100,
checkpoint_={"secret": SECRET},
metadata_={"secret": SECRET},
created_by=created_by,
)
@staticmethod
def _stage(
row_id: str,
*,
target_datasource_id: str,
promoted: bool,
) -> DatasourceStageRecord:
return DatasourceStageRecord(
id=row_id,
tenant_id="tenant-1",
target_datasource_id=target_datasource_id,
name=SECRET,
source_name=SECRET,
description=SECRET,
kind="table",
mode="static",
shape="tabular",
state="promoted" if promoted else "ready",
provider="upload",
provider_ref=SECRET,
schema_=[{"secret": SECRET}],
rows=[{"secret": SECRET}],
fingerprint="0" * 64,
row_count=1,
byte_count=100,
validation_={"secret": SECRET},
provenance_={"secret": SECRET},
metadata_={"secret": SECRET},
governance_={"secret": SECRET},
promoted_at=NOW if promoted else None,
promoted_materialization_id="materialization-1" if promoted else None,
created_by="account-1",
)
def test_canonical_selector_exports_only_minimized_attribution(self) -> None:
records = self.provider.search_subject(
self.session,
tenant_id="tenant-1",
subject=DsarSubjectRef(account_id="account-1"),
)
self.assertEqual(7, len(records))
self.assertEqual(
{"datasource_operator_attribution"},
{record.category for record in records},
)
exported = json.dumps([record.to_dict() for record in records])
self.assertNotIn(SECRET, exported)
self.assertNotIn("account-1", exported)
self.assertNotIn("datasource-other", exported)
def test_exact_datasource_returns_minimized_lifecycle_package(self) -> None:
records = self.provider.search_subject(
self.session,
tenant_id="tenant-1",
subject=DsarSubjectRef(
account_id="account-1",
external_references={
"datasources.datasource": "datasource:datasource-1"
},
),
)
self.assertEqual(7, len(records))
self.assertEqual(
{
"datasource",
"datasource_governance_reference",
"datasource_materialization",
"datasource_payload",
"datasource_stage",
"datasource_publication",
},
{record.resource_type for record in records},
)
exported = json.dumps([record.to_dict() for record in records])
self.assertNotIn(SECRET, exported)
self.assertNotIn("payload-orphan", exported)
actions = self.provider.plan_erasure(
self.session,
tenant_id="tenant-1",
subject=DsarSubjectRef(
external_references={"datasources.datasource": "datasource-1"}
),
records=records,
)
self.assertEqual({"manual_review"}, {action.kind for action in actions})
def test_exact_references_conflicts_and_tenants_fail_closed(self) -> None:
references = {
"datasources.governance_reference": "governance-1",
"datasources.materialization": "materialization-1",
"datasources.payload": "payload-referenced",
"datasources.stage": "stage-1",
"datasources.publication": "publication-1",
}
for key, value in references.items():
with self.subTest(key=key):
records = self.provider.search_subject(
self.session,
tenant_id="tenant-1",
subject=DsarSubjectRef(external_references={key: value}),
)
self.assertEqual(1, len(records))
mismatch = self.provider.search_subject(
self.session,
tenant_id="tenant-1",
subject=DsarSubjectRef(
account_id="account-2",
external_references={"datasources.stage": "stage-1"},
),
)
wrong_tenant = self.provider.search_subject(
self.session,
tenant_id="tenant-2",
subject=DsarSubjectRef(
external_references={"datasources.stage": "stage-1"}
),
)
conflict = self.provider.search_subject(
self.session,
tenant_id="tenant-1",
subject=DsarSubjectRef(
external_references={
"datasources.datasource": "datasource-1",
"datasources.catalogue": "different",
}
),
)
self.assertEqual((), mismatch)
self.assertEqual((), wrong_tenant)
self.assertEqual((), conflict)
def test_lifecycle_evidence_is_minimized_and_immutable(self) -> None:
evidence = DatasourceLifecycleEvidenceRecord(
id="evidence-1",
tenant_id="tenant-1",
subject_ref="datasource:datasource-1",
event_type="retention.applied",
occurred_at=NOW,
actor_ref="account-1",
policy_version="retention-v1",
policy_hash="1" * 64,
subject_digest="2" * 64,
event_hash="3" * 64,
details_={"secret": SECRET},
)
self.session.add(evidence)
self.session.commit()
exact_subject = DsarSubjectRef(
external_references={
"datasources.lifecycle_evidence": "evidence-1",
}
)
exact = self.provider.search_subject(
self.session,
tenant_id="tenant-1",
subject=exact_subject,
)
self.assertEqual(1, len(exact))
self.assertEqual("datasource_operator_attribution", exact[0].category)
self.assertNotIn(SECRET, json.dumps(exact[0].to_dict()))
actions = self.provider.plan_erasure(
self.session,
tenant_id="tenant-1",
subject=exact_subject,
records=exact,
)
self.assertEqual({"retain"}, {action.kind for action in actions})
canonical = self.provider.search_subject(
self.session,
tenant_id="tenant-1",
subject=DsarSubjectRef(account_id="account-1"),
)
self.assertIn(
"datasource_lifecycle_evidence",
{record.resource_type for record in canonical},
)
def test_transient_deletion_is_safe_and_idempotent(self) -> None:
subject = DsarSubjectRef(
external_references={
"datasources.stage": "stage-1",
"datasources.payload": "payload-orphan",
}
)
records = self.provider.search_subject(
self.session,
tenant_id="tenant-1",
subject=subject,
)
actions = self.provider.plan_erasure(
self.session,
tenant_id="tenant-1",
subject=subject,
records=records,
)
self.assertEqual({"delete"}, {action.kind for action in actions})
first = self.provider.execute_erasure(
self.session,
tenant_id="tenant-1",
subject=subject,
actions=actions,
request_id="dsar-1",
)
second = self.provider.execute_erasure(
self.session,
tenant_id="tenant-1",
subject=subject,
actions=actions,
request_id="dsar-1-retry",
)
self.assertTrue(all(result.status == "executed" for result in first))
self.assertTrue(all(result.status == "unchanged" for result in second))
self.assertIsNone(self.session.get(DatasourceStageRecord, "stage-1"))
self.assertIsNone(self.session.get(DatasourcePayloadRecord, "payload-orphan"))
self.assertIsNotNone(
self.session.get(DatasourcePayloadRecord, "payload-referenced")
)
def test_published_and_attribution_state_requires_review_or_retention(self) -> None:
direct_subject = DsarSubjectRef(
external_references={
"datasources.materialization": "materialization-1",
"datasources.payload": "payload-referenced",
"datasources.stage": "stage-promoted",
}
)
direct = self.provider.search_subject(
self.session,
tenant_id="tenant-1",
subject=direct_subject,
)
direct_actions = self.provider.plan_erasure(
self.session,
tenant_id="tenant-1",
subject=direct_subject,
records=direct,
)
self.assertEqual(
{"manual_review"},
{action.kind for action in direct_actions},
)
canonical_subject = DsarSubjectRef(account_id="account-1")
canonical = self.provider.search_subject(
self.session,
tenant_id="tenant-1",
subject=canonical_subject,
)
canonical_actions = self.provider.plan_erasure(
self.session,
tenant_id="tenant-1",
subject=canonical_subject,
records=canonical,
)
self.assertEqual({"retain"}, {action.kind for action in canonical_actions})
results = self.provider.execute_erasure(
self.session,
tenant_id="tenant-1",
subject=canonical_subject,
actions=canonical_actions,
request_id="dsar-2",
)
self.assertTrue(all(result.status == "blocked" for result in results))
def test_foreign_records_and_actions_are_rejected(self) -> None:
subject = DsarSubjectRef(account_id="account-1")
with self.assertRaisesRegex(ValueError, "foreign provider record"):
self.provider.plan_erasure(
self.session,
tenant_id="tenant-1",
subject=subject,
records=(
DsarRecordRef(
provider_id="cases",
module_id="cases",
resource_type="case",
resource_id="case-1",
category="case",
title="Case",
),
),
)
with self.assertRaisesRegex(ValueError, "foreign provider action"):
self.provider.execute_erasure(
self.session,
tenant_id="tenant-1",
subject=subject,
actions=(
DsarErasureActionRef(
action_id="cases:delete:case:case-1",
provider_id="cases",
module_id="cases",
kind="delete",
resource_type="case",
resource_id="case-1",
title="Delete case",
rationale="Foreign",
executable=True,
),
),
request_id="dsar-3",
)
def test_core_workflow_reports_active_and_inactive_provider(self) -> None:
row = create_data_subject_request(
self.session,
tenant_id="tenant-1",
reference="DSAR-DATASOURCES-1",
request_kind="access_and_erasure",
subject=DsarSubjectRef(account_id="account-1"),
purpose="Respond to a verified request.",
legal_basis="Article 15 and 17 GDPR",
due_at=None,
requested_by_account_id="privacy-officer",
)
self.session.commit()
search_data_subject_request(
self.session,
registry=_Registry(self.provider),
row=row,
expected_revision=1,
)
self.assertEqual(
[DATASOURCES_DSAR_CAPABILITY],
row.coverage["provider_capabilities"],
)
self.assertEqual(7, row.search_result["record_count"])
inactive = create_data_subject_request(
self.session,
tenant_id="tenant-1",
reference="DSAR-DATASOURCES-2",
request_kind="access",
subject=DsarSubjectRef(account_id="account-1"),
purpose="Respond to a verified request.",
legal_basis="Article 15 GDPR",
due_at=None,
requested_by_account_id="privacy-officer",
)
self.session.commit()
search_data_subject_request(
self.session,
registry=_Registry(self.provider, active=False),
row=inactive,
expected_revision=1,
)
self.assertEqual([], inactive.coverage["provider_capabilities"])
self.assertEqual(
[DATASOURCES_DSAR_CAPABILITY],
inactive.coverage["inactive_provider_capabilities"],
)
self.assertEqual(0, inactive.search_result["record_count"])
def test_manifest_registers_and_documents_capability(self) -> None:
self.assertIn(DATASOURCES_DSAR_CAPABILITY, manifest.capability_factories)
self.assertIn(
DATASOURCES_DSAR_CAPABILITY,
manifest.capability_documentation,
)
self.assertIn(
DATASOURCES_DSAR_CAPABILITY,
{item.name for item in manifest.provides_interfaces},
)
self.assertTrue(
any(
topic.id == "datasources.data-subject-requests"
and {"admin", "user"}.issubset(topic.documentation_types)
for topic in manifest.documentation
)
)
if __name__ == "__main__":
unittest.main()
@@ -18,6 +18,8 @@ class DatasourcesInterfaceDocumentationContractTests(unittest.TestCase):
"datasources.origins", "datasources.origins",
"datasources.governance", "datasources.governance",
"datasources.preview", "datasources.preview",
"datasources.lifecycle-evidence",
"datasources.retention",
}, },
{item.id for item in frontend.view_surfaces}, # type: ignore[union-attr] {item.id for item in frontend.view_surfaces}, # type: ignore[union-attr]
) )
@@ -27,16 +29,26 @@ class DatasourcesInterfaceDocumentationContractTests(unittest.TestCase):
lifecycle = topics["datasources.lifecycle"] lifecycle = topics["datasources.lifecycle"]
governance = topics["datasources.governance"] governance = topics["datasources.governance"]
quality = topics["datasources.quality-gates"] quality = topics["datasources.quality-gates"]
visibility = topics["datasources.visibility"]
reference = topics["datasources.reference.fields-and-consequences"] reference = topics["datasources.reference.fields-and-consequences"]
lifecycle_controls = topics["datasources.approval-and-retention"]
self.assertIn("datasources.staging", lifecycle.metadata["help_contexts"]) self.assertIn("datasources.staging", lifecycle.metadata["help_contexts"])
self.assertIn("datasources.field.authority-mode", governance.metadata["help_contexts"]) self.assertIn("datasources.field.authority-mode", governance.metadata["help_contexts"])
self.assertIn("datasources.staging.validation", quality.metadata["help_contexts"]) self.assertIn("datasources.staging.validation", quality.metadata["help_contexts"])
self.assertIn(
"datasources.field.visibility-policy",
visibility.metadata["help_contexts"],
)
self.assertIn("fails closed", visibility.body)
self.assertIn("policy version and hash", quality.body) self.assertIn("policy version and hash", quality.body)
self.assertTrue(quality.metadata["limitations"]) self.assertTrue(quality.metadata["limitations"])
self.assertIn("datasources.action.promote", reference.metadata["help_contexts"]) self.assertIn("datasources.action.promote", reference.metadata["help_contexts"])
self.assertIn("freeze", reference.metadata["consequence_classes"]) self.assertIn("freeze", reference.metadata["consequence_classes"])
self.assertIn("retire", reference.metadata["consequence_classes"]) self.assertIn("retire", reference.metadata["consequence_classes"])
self.assertIn("datasources.action.approve", lifecycle_controls.metadata["help_contexts"])
self.assertIn("hash-chained lifecycle evidence", lifecycle_controls.body)
self.assertTrue(lifecycle_controls.translations["de"]["body"])
if __name__ == "__main__": if __name__ == "__main__":
+198
View File
@@ -11,6 +11,7 @@ from govoplan_core.core.change_sequence import ChangeSequenceEntry
from govoplan_core.core.datasources import ( from govoplan_core.core.datasources import (
CAPABILITY_DATASOURCE_ORIGINS, CAPABILITY_DATASOURCE_ORIGINS,
DatasourceAccessError, DatasourceAccessError,
DatasourceArtifactReference,
DatasourceField, DatasourceField,
DatasourceGovernance, DatasourceGovernance,
DatasourceOrigin, DatasourceOrigin,
@@ -30,6 +31,7 @@ from govoplan_core.core.tabular_sources import (
from govoplan_core.db.base import Base, utcnow from govoplan_core.db.base import Base, utcnow
from govoplan_datasources.backend.db.models import ( from govoplan_datasources.backend.db.models import (
DatasourceGovernanceReferenceRecord, DatasourceGovernanceReferenceRecord,
DatasourceLifecycleEvidenceRecord,
DatasourceMaterializationRecord, DatasourceMaterializationRecord,
DatasourcePayloadRecord, DatasourcePayloadRecord,
DatasourcePayloadRowRecord, DatasourcePayloadRowRecord,
@@ -44,6 +46,7 @@ from govoplan_datasources.backend.service import (
SqlDatasourceProvider, SqlDatasourceProvider,
) )
from govoplan_datasources.backend.payloads import ( from govoplan_datasources.backend.payloads import (
ExternalArtifactPayloadBackend,
create_database_rows_payload, create_database_rows_payload,
finalize_payload_deletion, finalize_payload_deletion,
mark_unreferenced_payload_for_deletion, mark_unreferenced_payload_for_deletion,
@@ -166,6 +169,56 @@ class FakeRegistry:
return self.origin_provider return self.origin_provider
class FakeArtifactBackend:
backend = "test_artifact"
def __init__(self) -> None:
self.verified: list[str] = []
self.deleted: list[str] = []
def read_rows(
self,
_session,
*,
tenant_id: str,
artifact: DatasourceArtifactReference,
offset: int,
limit: int,
):
self.assert_tenant(tenant_id)
stop = min(artifact.row_count, offset + limit)
return tuple(
{"id": index, "result": "match"}
for index in range(offset, stop)
)
def verify(
self,
_session,
*,
tenant_id: str,
artifact: DatasourceArtifactReference,
) -> None:
self.assert_tenant(tenant_id)
if not artifact.locator.startswith("artifact:"):
raise DatasourceUnavailableError("Unknown test artifact.")
self.verified.append(artifact.locator)
def delete(
self,
_session,
*,
tenant_id: str,
artifact: DatasourceArtifactReference,
) -> None:
self.assert_tenant(tenant_id)
self.deleted.append(artifact.locator)
def assert_tenant(self, tenant_id: str) -> None:
if tenant_id != "tenant-1":
raise AssertionError("Artifact backend crossed a tenant boundary.")
class DatasourceLifecycleTests(unittest.TestCase): class DatasourceLifecycleTests(unittest.TestCase):
def setUp(self) -> None: def setUp(self) -> None:
self.engine = create_engine("sqlite:///:memory:") self.engine = create_engine("sqlite:///:memory:")
@@ -174,6 +227,7 @@ class DatasourceLifecycleTests(unittest.TestCase):
tables=[ tables=[
DatasourceRecord.__table__, DatasourceRecord.__table__,
DatasourceGovernanceReferenceRecord.__table__, DatasourceGovernanceReferenceRecord.__table__,
DatasourceLifecycleEvidenceRecord.__table__,
DatasourcePayloadRecord.__table__, DatasourcePayloadRecord.__table__,
DatasourcePayloadRowRecord.__table__, DatasourcePayloadRowRecord.__table__,
DatasourceMaterializationRecord.__table__, DatasourceMaterializationRecord.__table__,
@@ -185,8 +239,10 @@ class DatasourceLifecycleTests(unittest.TestCase):
self.Session = sessionmaker(bind=self.engine) self.Session = sessionmaker(bind=self.engine)
self.session = self.Session() self.session = self.Session()
self.origins = FakeOriginProvider() self.origins = FakeOriginProvider()
self.artifacts = FakeArtifactBackend()
self.provider = SqlDatasourceProvider( self.provider = SqlDatasourceProvider(
registry=FakeRegistry(self.origins), registry=FakeRegistry(self.origins),
payload_backends=(ExternalArtifactPayloadBackend(self.artifacts),),
) )
def tearDown(self) -> None: def tearDown(self) -> None:
@@ -194,6 +250,7 @@ class DatasourceLifecycleTests(unittest.TestCase):
Base.metadata.drop_all( Base.metadata.drop_all(
self.engine, self.engine,
tables=[ tables=[
DatasourceLifecycleEvidenceRecord.__table__,
DatasourceStageRecord.__table__, DatasourceStageRecord.__table__,
DatasourcePublicationRecord.__table__, DatasourcePublicationRecord.__table__,
DatasourceMaterializationRecord.__table__, DatasourceMaterializationRecord.__table__,
@@ -667,6 +724,147 @@ class DatasourceLifecycleTests(unittest.TestCase):
self.session.query(DatasourcePublicationRecord).count(), self.session.query(DatasourcePublicationRecord).count(),
) )
def test_artifact_publication_pins_large_payload_and_supports_bounded_reads(
self,
) -> None:
artifact = DatasourceArtifactReference(
backend="test_artifact",
locator="artifact:monthly-output",
checksum="a" * 64,
row_count=25_000,
byte_count=12_000_000,
schema=(
DatasourceField("id", "integer", nullable=False),
DatasourceField("result", "string", nullable=False),
),
fingerprint="b" * 64,
)
published = self.provider.publish_rows(
self.session,
principal(scopes=(SOURCE_WRITE_SCOPE, CATALOGUE_READ_SCOPE)),
request=DatasourcePublicationRequest(
producer_module="dataflow",
producer_run_ref="dataflow-run:large-output",
idempotency_key="large-output",
name="Large output",
source_name="large_output",
artifact=artifact,
),
)
preview = self.provider.read_datasource(
self.session,
principal(scopes=(CATALOGUE_READ_SCOPE,)),
request=DatasourceReadRequest(
datasource_ref=published.datasource.ref,
offset=10,
limit=3,
),
)
self.assertEqual("published", published.status)
self.assertEqual(25_000, published.materialization.row_count)
self.assertEqual(
[{"id": 10, "result": "match"},
{"id": 11, "result": "match"},
{"id": 12, "result": "match"}],
list(preview.rows),
)
self.assertEqual(
["artifact:monthly-output", "artifact:monthly-output"],
self.artifacts.verified,
)
def test_unattested_artifact_quality_rules_require_review_without_becoming_current(
self,
) -> None:
published = self.provider.publish_rows(
self.session,
principal(scopes=(SOURCE_WRITE_SCOPE,)),
request=DatasourcePublicationRequest(
producer_module="reporting",
producer_run_ref="report-run:review",
idempotency_key="review-output",
name="Review output",
source_name="review_output",
artifact=DatasourceArtifactReference(
backend="test_artifact",
locator="artifact:review-output",
checksum="c" * 64,
row_count=2,
byte_count=128,
schema=(DatasourceField("id", "integer", False),),
fingerprint="d" * 64,
),
governance=DatasourceGovernance(
quality_policy={
"version": "unique-id-v1",
"rules": [
{
"id": "unique-id",
"type": "unique",
"fields": ["id"],
}
],
}
),
),
)
record = self.session.get(
DatasourceRecord,
published.datasource.ref.removeprefix("datasource:"),
)
self.assertEqual("review_required", published.status)
self.assertEqual("review_required", published.materialization.state)
self.assertIsNotNone(record)
assert record is not None
self.assertIsNone(record.current_materialization_id)
def test_artifact_warning_is_a_notification_ready_terminal_state(self) -> None:
published = self.provider.publish_rows(
self.session,
principal(scopes=(SOURCE_WRITE_SCOPE,)),
request=DatasourcePublicationRequest(
producer_module="dataflow",
producer_run_ref="dataflow-run:warning",
idempotency_key="warning-output",
name="Warning output",
source_name="warning_output",
artifact=DatasourceArtifactReference(
backend="test_artifact",
locator="artifact:warning-output",
checksum="e" * 64,
row_count=1,
byte_count=64,
schema=(DatasourceField("id", "integer", False),),
fingerprint="f" * 64,
validation={
"status": "warning",
"warnings": [
{
"severity": "warning",
"code": "producer.partial_match",
"message": "One source used a fallback match.",
}
],
},
),
),
)
record = self.session.get(
DatasourcePublicationRecord,
published.ref.removeprefix("publication:"),
)
self.assertEqual("published_with_warnings", published.status)
self.assertIsNotNone(record)
assert record is not None
self.assertEqual(
"published_with_warnings",
record.status,
)
def test_publication_idempotency_key_rejects_different_output(self) -> None: def test_publication_idempotency_key_rejects_different_output(self) -> None:
producer = principal(scopes=(SOURCE_WRITE_SCOPE,)) producer = principal(scopes=(SOURCE_WRITE_SCOPE,))
base = DatasourcePublicationRequest( base = DatasourcePublicationRequest(
+497
View File
@@ -0,0 +1,497 @@
from __future__ import annotations
import unittest
from datetime import timedelta
from sqlalchemy import create_engine, select
from sqlalchemy.orm import sessionmaker
from govoplan_core.auth import ApiPrincipal
from govoplan_core.core.access import PrincipalRef
from govoplan_core.core.datasources import (
CAPABILITY_DATASOURCE_ORIGINS,
DatasourceField,
DatasourceGovernance,
DatasourceOrigin,
DatasourceOriginReadRequest,
DatasourceOriginReadResult,
DatasourceReadRequest,
DatasourceStageInput,
DatasourceUnavailableError,
DatasourceValidationError,
)
from govoplan_core.core.tabular_sources import TabularPushdown, TabularSourceHealth
from govoplan_core.db.base import Base
from govoplan_datasources.backend.db.models import (
DatasourceGovernanceReferenceRecord,
DatasourceLifecycleEvidenceRecord,
DatasourceMaterializationRecord,
DatasourcePayloadRecord,
DatasourcePayloadRowRecord,
DatasourcePublicationRecord,
DatasourceRecord,
DatasourceStageRecord,
)
from govoplan_datasources.backend.service import (
CATALOGUE_READ_SCOPE,
SOURCE_WRITE_SCOPE,
STAGE_APPROVE_SCOPE,
STAGE_WRITE_SCOPE,
SqlDatasourceProvider,
)
def principal(
account_id: str,
*,
scopes: tuple[str, ...],
) -> ApiPrincipal:
return ApiPrincipal(
principal=PrincipalRef(
account_id=account_id,
membership_id=f"membership-{account_id}",
tenant_id="tenant-1",
scopes=frozenset(scopes),
),
account=object(),
user=object(),
)
class _OriginProvider:
def __init__(self) -> None:
self.rows = [{"id": 1, "status": "initial"}]
def origin(self) -> DatasourceOrigin:
return DatasourceOrigin(
ref="origin:cases",
source_name="cached_cases",
name="Cached cases",
kind="database",
shape="tabular",
supported_modes=("cached",),
provider="connectors.test",
schema=(
DatasourceField("id", "integer", False),
DatasourceField("status", "string", False),
),
schema_version="1",
fingerprint=f"cases-{self.rows[-1]['status']}",
row_count=len(self.rows),
source_mode="cached",
pushdown=TabularPushdown(pagination=True),
health=TabularSourceHealth(
status="healthy",
code="origin.ready",
summary="Origin is ready.",
),
)
def list_origins(self, _session, _principal, *, query="", limit=100):
del limit
origin = self.origin()
return (origin,) if query.casefold() in origin.name.casefold() else ()
def get_origin(self, _session, _principal, *, origin_ref):
return self.origin() if origin_ref == "origin:cases" else None
def read_origin(
self,
_session,
_principal,
*,
request: DatasourceOriginReadRequest,
) -> DatasourceOriginReadResult:
rows = self.rows[request.offset : request.offset + request.limit]
return DatasourceOriginReadResult(
origin=self.origin(),
rows=tuple(dict(row) for row in rows),
total_rows=len(self.rows),
truncated=request.offset + len(rows) < len(self.rows),
returned_bytes=64,
elapsed_ms=1,
effective_row_limit=request.limit,
effective_byte_limit=request.max_bytes,
effective_timeout_ms=request.timeout_ms,
)
class _Registry:
def __init__(self, origins: _OriginProvider) -> None:
self.origins = origins
def has_capability(self, name: str) -> bool:
return name == CAPABILITY_DATASOURCE_ORIGINS
def capability(self, name: str):
if not self.has_capability(name):
raise KeyError(name)
return self.origins
class DatasourceLifecycleGovernanceTests(unittest.TestCase):
def setUp(self) -> None:
self.engine = create_engine("sqlite:///:memory:")
self.tables = [
DatasourceRecord.__table__,
DatasourceGovernanceReferenceRecord.__table__,
DatasourcePayloadRecord.__table__,
DatasourcePayloadRowRecord.__table__,
DatasourceMaterializationRecord.__table__,
DatasourceStageRecord.__table__,
DatasourcePublicationRecord.__table__,
DatasourceLifecycleEvidenceRecord.__table__,
]
Base.metadata.create_all(self.engine, tables=self.tables)
self.session = sessionmaker(bind=self.engine)()
self.origins = _OriginProvider()
self.provider = SqlDatasourceProvider(registry=_Registry(self.origins))
self.writer = principal(
"writer",
scopes=(
CATALOGUE_READ_SCOPE,
SOURCE_WRITE_SCOPE,
STAGE_WRITE_SCOPE,
),
)
def tearDown(self) -> None:
self.session.close()
Base.metadata.drop_all(self.engine, tables=list(reversed(self.tables)))
self.engine.dispose()
def test_approval_quorum_is_attributable_idempotent_and_separated(self) -> None:
stage = self.provider.create_stage(
self.session,
self.writer,
stage=DatasourceStageInput(
name="Approval register",
source_name="approval_register",
kind="upload",
mode="static",
shape="tabular",
rows=({"id": 1},),
governance=DatasourceGovernance(
approval_policy={
"version": "promotion-v2",
"required": True,
"required_approvals": 2,
"separation_of_duties": True,
}
),
),
)
self.assertEqual("awaiting_approval", stage.state)
policy_hash = str(stage.approval["policy_hash"])
subject_digest = str(stage.approval["subject_digest"])
creator_approver = principal(
"writer",
scopes=(STAGE_APPROVE_SCOPE,),
)
with self.assertRaisesRegex(
DatasourceValidationError,
"creator cannot approve",
):
self.provider.decide_stage(
self.session,
creator_approver,
stage_ref=stage.ref,
decision="approve",
reason="Creator review",
expected_policy_hash=policy_hash,
expected_subject_digest=subject_digest,
)
first, first_hash, replayed = self.provider.decide_stage(
self.session,
principal("approver-1", scopes=(STAGE_APPROVE_SCOPE,)),
stage_ref=stage.ref,
decision="approve",
reason="Quality and schema evidence reviewed.",
expected_policy_hash=policy_hash,
expected_subject_digest=subject_digest,
)
self.assertFalse(replayed)
self.assertTrue(first_hash)
self.assertEqual("awaiting_approval", first.state)
replay, replay_hash, replayed = self.provider.decide_stage(
self.session,
principal("approver-1", scopes=(STAGE_APPROVE_SCOPE,)),
stage_ref=stage.ref,
decision="approve",
reason="Quality and schema evidence reviewed.",
expected_policy_hash=policy_hash,
expected_subject_digest=subject_digest,
)
self.assertTrue(replayed)
self.assertIsNone(replay_hash)
self.assertEqual(1, replay.approval["approval_count"])
approved, second_hash, replayed = self.provider.decide_stage(
self.session,
principal("approver-2", scopes=(STAGE_APPROVE_SCOPE,)),
stage_ref=stage.ref,
decision="approve",
reason="Authority and promotion impact reviewed.",
expected_policy_hash=policy_hash,
expected_subject_digest=subject_digest,
)
self.assertFalse(replayed)
self.assertTrue(second_hash)
self.assertEqual("ready", approved.state)
datasource, materialization = self.provider.promote_stage(
self.session,
self.writer,
stage_ref=stage.ref,
)
self.assertEqual(2, materialization.provenance["stage_approval"]["approval_count"])
self.assertTrue(materialization.provenance["promotion_evidence_hash"])
self.assertEqual(materialization.ref, datasource.current_materialization_ref)
evidence = tuple(
reversed(
self.provider.list_lifecycle_evidence(
self.session,
self.writer,
subject_ref=stage.ref,
)
)
)
self.assertEqual(
["stage.validated", "stage.approved", "stage.approved", "stage.promoted"],
[item.event_type for item in evidence],
)
for previous, current in zip(evidence, evidence[1:], strict=False):
self.assertEqual(previous.event_hash, current.previous_event_hash)
def test_approved_refresh_stage_is_required_before_current_changes(self) -> None:
datasource = self.provider.register_origin(
self.session,
self.writer,
origin_ref="origin:cases",
name="Cached cases",
source_name="cached_cases",
mode="cached",
governance=DatasourceGovernance(
approval_policy={
"version": "refresh-v1",
"required": True,
"required_approvals": 1,
}
),
)
original_ref = datasource.current_materialization_ref
self.origins.rows = [{"id": 1, "status": "refreshed"}]
with self.assertRaisesRegex(
DatasourceValidationError,
"approved refresh stage",
):
self.provider.refresh_datasource(
self.session,
self.writer,
datasource_ref=datasource.ref,
)
stage = self.provider.prepare_refresh(
self.session,
self.writer,
datasource_ref=datasource.ref,
)
self.assertEqual("awaiting_approval", stage.state)
approved, _, _ = self.provider.decide_stage(
self.session,
principal("refresh-approver", scopes=(STAGE_APPROVE_SCOPE,)),
stage_ref=stage.ref,
decision="approve",
reason="The connector delta and schema are acceptable.",
expected_policy_hash=str(stage.approval["policy_hash"]),
expected_subject_digest=str(stage.approval["subject_digest"]),
)
self.assertEqual("ready", approved.state)
refreshed, current = self.provider.promote_stage(
self.session,
self.writer,
stage_ref=stage.ref,
)
self.assertNotEqual(original_ref, refreshed.current_materialization_ref)
self.assertEqual(current.ref, refreshed.current_materialization_ref)
preview = self.provider.read_datasource(
self.session,
self.writer,
request=DatasourceReadRequest(datasource_ref=datasource.ref),
)
self.assertEqual("refreshed", preview.rows[0]["status"])
def test_retention_preview_blocks_current_and_held_evidence_then_purges_payload(self) -> None:
governance = DatasourceGovernance(
retention_policy={
"version": "records-v3",
"enabled": True,
"materialization_days": 1,
"frozen_evidence_days": 1,
}
)
first_stage = self.provider.create_stage(
self.session,
self.writer,
stage=DatasourceStageInput(
name="Retained register",
source_name="retained_register",
kind="upload",
mode="static",
shape="tabular",
rows=({"id": 1},),
governance=governance,
),
)
datasource, first = self.provider.promote_stage(
self.session,
self.writer,
stage_ref=first_stage.ref,
)
second_stage = self.provider.create_stage(
self.session,
self.writer,
stage=DatasourceStageInput(
name="Retained register",
source_name="retained_register",
kind="upload",
mode="static",
shape="tabular",
rows=({"id": 2},),
target_datasource_ref=datasource.ref,
),
)
_, second = self.provider.promote_stage(
self.session,
self.writer,
stage_ref=second_stage.ref,
)
as_of = max(first.created_at, second.created_at) + timedelta(days=2)
plan = self.provider.preview_retention(
self.session,
principal("admin", scopes=("datasources:source:admin",)),
as_of=as_of,
)
by_ref = {item.ref: item for item in plan.candidates}
self.assertTrue(by_ref[first.ref].eligible)
self.assertFalse(by_ref[second.ref].eligible)
self.assertIn("current_materialization", by_ref[second.ref].blockers)
first_row = self.session.scalar(
select(DatasourceMaterializationRecord).where(
DatasourceMaterializationRecord.id
== first.ref.removeprefix("materialization:")
)
)
payload_id = first_row.payload_id
disposed, evidence_hashes = self.provider.apply_retention(
self.session,
principal("admin", scopes=("datasources:source:admin",)),
as_of=as_of,
plan_hash=plan.plan_hash,
target_refs=(first.ref,),
)
self.assertEqual((first.ref,), disposed)
self.assertEqual(1, len(evidence_hashes))
self.assertIsNone(self.session.get(DatasourcePayloadRecord, payload_id))
self.assertIsNotNone(first_row.disposed_at)
self.assertEqual("disposed", first_row.state)
with self.assertRaises(DatasourceUnavailableError):
self.provider.read_datasource(
self.session,
self.writer,
request=DatasourceReadRequest(
datasource_ref=datasource.ref,
materialization_ref=first.ref,
),
)
held = self.provider.update_datasource_governance(
self.session,
self.writer,
datasource_ref=datasource.ref,
governance=DatasourceGovernance(
retention_policy=governance.retention_policy,
hold_refs=("hold:legal-1",),
),
)
frozen = self.provider.freeze_datasource(
self.session,
self.writer,
datasource_ref=held.ref,
label="Legal evidence",
)
held_plan = self.provider.preview_retention(
self.session,
principal("admin", scopes=("datasources:source:admin",)),
as_of=frozen.created_at + timedelta(days=2),
)
held_candidate = next(item for item in held_plan.candidates if item.ref == frozen.ref)
self.assertFalse(held_candidate.eligible)
self.assertIn("legal_hold", held_candidate.blockers)
def test_retention_deletes_only_an_explicitly_selected_eligible_stage(self) -> None:
stage = self.provider.create_stage(
self.session,
self.writer,
stage=DatasourceStageInput(
name="Transient import",
source_name="transient_import",
kind="upload",
mode="static",
shape="tabular",
rows=({"id": 1},),
governance=DatasourceGovernance(
retention_policy={
"version": "stage-retention-v1",
"enabled": True,
"stage_days": 1,
}
),
),
)
admin = principal(
"admin",
scopes=("datasources:source:admin",),
)
as_of = stage.created_at + timedelta(days=2)
plan = self.provider.preview_retention(
self.session,
admin,
as_of=as_of,
)
candidate = next(item for item in plan.candidates if item.ref == stage.ref)
self.assertTrue(candidate.eligible)
self.assertEqual("delete_stage", candidate.disposition)
disposed, evidence_hashes = self.provider.apply_retention(
self.session,
admin,
as_of=as_of,
plan_hash=plan.plan_hash,
target_refs=(stage.ref,),
)
self.assertEqual((stage.ref,), disposed)
self.assertEqual(1, len(evidence_hashes))
self.assertIsNone(
self.session.get(
DatasourceStageRecord,
stage.ref.removeprefix("stage:"),
)
)
evidence = self.provider.list_lifecycle_evidence(
self.session,
self.writer,
subject_ref=stage.ref,
)
self.assertEqual("retention.stage_deleted", evidence[0].event_type)
self.assertEqual(plan.plan_hash, evidence[0].details_["plan_hash"])
if __name__ == "__main__":
unittest.main()
+37
View File
@@ -0,0 +1,37 @@
from __future__ import annotations
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
PAGE = ROOT / "webui" / "src" / "features" / "datasources" / "DatasourcesPage.tsx"
API = ROOT / "webui" / "src" / "api" / "datasources.ts"
class DatasourceLifecycleGovernanceUiTests(unittest.TestCase):
def test_stage_decisions_and_retention_use_shared_semantic_primitives(self) -> None:
source = PAGE.read_text(encoding="utf-8")
self.assertIn("<WorkspaceActionBar", source)
self.assertIn("function StageDecisionDialog", source)
self.assertIn("function RetentionDialog", source)
self.assertIn("<ConfirmDialog", source)
self.assertIn('variant="danger"', source)
self.assertIn("datasources.field.approval-policy", source)
self.assertIn("datasources.field.retention-policy-contract", source)
self.assertIn('state === "awaiting_approval"', source)
def test_client_binds_decisions_and_apply_to_server_evidence(self) -> None:
source = API.read_text(encoding="utf-8")
self.assertIn("expected_policy_hash", source)
self.assertIn("expected_subject_digest", source)
self.assertIn("/lifecycle-evidence", source)
self.assertIn("/retention/plan", source)
self.assertIn("plan_hash: plan.plan_hash", source)
self.assertIn("/refresh/stage", source)
if __name__ == "__main__":
unittest.main()
+4
View File
@@ -31,6 +31,10 @@ class DatasourceManifestTests(unittest.TestCase):
if item.name == "connectors.datasource_origins" if item.name == "connectors.datasource_origins"
) )
self.assertTrue(requirement.optional) self.assertTrue(requirement.optional)
self.assertIn(
"datasources:stage:approve",
{item.scope for item in manifest.permissions},
)
if __name__ == "__main__": if __name__ == "__main__":
+24 -1
View File
@@ -34,7 +34,7 @@ class DatasourceMigrationTests(unittest.TestCase):
try: try:
with engine.connect() as connection: with engine.connect() as connection:
self.assertIn( self.assertIn(
"b8d2f5a0c3e7", "d1a7c3e9f5b2",
set(MigrationContext.configure(connection).get_current_heads()), set(MigrationContext.configure(connection).get_current_heads()),
) )
catalogue_columns = { catalogue_columns = {
@@ -50,13 +50,36 @@ class DatasourceMigrationTests(unittest.TestCase):
"publication_state", "publication_state",
"owner_ref", "owner_ref",
"quality_policy", "quality_policy",
"access_policy_ref",
"visibility_policy",
"approval_policy",
"retention_policy",
"dependency_refs", "dependency_refs",
}.issubset(catalogue_columns) }.issubset(catalogue_columns)
) )
stage_columns = {
item["name"]
for item in inspect(connection).get_columns(
"datasource_stages"
)
}
materialization_columns = {
item["name"]
for item in inspect(connection).get_columns(
"datasource_materializations"
)
}
self.assertIn("approval", stage_columns)
self.assertTrue(
{"disposed_at", "disposition"}.issubset(
materialization_columns
)
)
self.assertEqual( self.assertEqual(
{ {
"datasource_catalogue", "datasource_catalogue",
"datasource_governance_references", "datasource_governance_references",
"datasource_lifecycle_evidence",
"datasource_materializations", "datasource_materializations",
"datasource_payload_rows", "datasource_payload_rows",
"datasource_payloads", "datasource_payloads",
+162
View File
@@ -0,0 +1,162 @@
from __future__ import annotations
from types import SimpleNamespace
import unittest
from sqlalchemy import create_engine
from sqlalchemy.orm import Session
from govoplan_core.auth import ApiPrincipal
from govoplan_core.core.access import PrincipalRef
from govoplan_core.core.search import (
SearchAuthorizationRequest,
SearchBackfillRequest,
SearchResourceReference,
)
from govoplan_core.db.base import Base
from govoplan_datasources.backend.db.models import DatasourceRecord
from govoplan_datasources.backend.manifest import get_manifest
from govoplan_datasources.backend.search_source import (
DatasourcesSearchSource,
PROVIDER_ID,
RESOURCE_TYPE,
)
from govoplan_datasources.backend.service import ADMIN_SCOPE, CATALOGUE_READ_SCOPE
class DatasourcesSearchSourceTests(unittest.TestCase):
def setUp(self) -> None:
self.engine = create_engine("sqlite://")
Base.metadata.create_all(
self.engine,
tables=(DatasourceRecord.__table__,),
)
self.session = Session(self.engine)
self.session.add_all(
(
_datasource("source-1", "tenant-1", "Monthly source"),
_datasource("source-other", "tenant-2", "Other tenant"),
)
)
self.session.commit()
self.source = DatasourcesSearchSource()
def tearDown(self) -> None:
self.session.close()
self.engine.dispose()
def test_manifest_registers_optional_search_source(self) -> None:
manifest = get_manifest()
self.assertIn("search", manifest.optional_dependencies)
self.assertIn(
PROVIDER_ID,
{registration.id for registration in manifest.search_sources},
)
def test_backfill_is_tenant_bound_and_excludes_protected_payloads(self) -> None:
page = self.source.backfill(
self.session,
request=SearchBackfillRequest(
tenant_id="tenant-1",
provider_id=PROVIDER_ID,
resource_type=RESOURCE_TYPE,
rebuild_id="rebuild-1",
),
)
self.assertEqual(("source-1",), tuple(doc.resource_id for doc in page.documents))
document = page.documents[0]
serialized = repr(document)
self.assertNotIn("must-not-be-indexed", serialized)
self.assertNotIn("secret_field", serialized)
self.assertNotIn("credential-1", serialized)
self.assertTrue(document.requires_authorization_recheck)
self.assertEqual(
"/datasources?datasource=datasource%3Asource-1",
document.url,
)
def test_authorization_rechecks_scope_tenant_and_current_existence(self) -> None:
reference = SearchResourceReference(
tenant_id="tenant-1",
module_id="datasources",
resource_type=RESOURCE_TYPE,
resource_id="source-1",
)
request = SearchAuthorizationRequest(reference=reference, source_revision="1")
self.assertTrue(
self.source.authorize(
self.session,
_principal({CATALOGUE_READ_SCOPE}),
requests=(request,),
)[reference.key]
)
self.assertTrue(
self.source.authorize(
self.session,
_principal({ADMIN_SCOPE}),
requests=(request,),
)[reference.key]
)
self.assertFalse(
self.source.authorize(
self.session,
_principal(set()),
requests=(request,),
)[reference.key]
)
other_reference = SearchResourceReference(
tenant_id="tenant-2",
module_id="datasources",
resource_type=RESOURCE_TYPE,
resource_id="source-other",
)
other_request = SearchAuthorizationRequest(
reference=other_reference,
source_revision="1",
)
self.assertFalse(
self.source.authorize(
self.session,
_principal({CATALOGUE_READ_SCOPE}),
requests=(other_request,),
)[other_reference.key]
)
def _datasource(identifier: str, tenant_id: str, name: str) -> DatasourceRecord:
return DatasourceRecord(
id=identifier,
tenant_id=tenant_id,
source_name=identifier.replace("-", "_"),
name=name,
description="Safe catalogue description",
kind="database",
mode="cached",
shape="tabular",
status="active",
provider="connectors.sql",
provider_ref="credential-1",
schema_=[{"name": "secret_field", "type": "string"}],
provenance_={"query": "must-not-be-indexed"},
metadata_={"password": "must-not-be-indexed"},
)
def _principal(scopes: set[str]) -> ApiPrincipal:
return ApiPrincipal(
principal=PrincipalRef(
account_id="account-1",
membership_id="user-1",
tenant_id="tenant-1",
scopes=frozenset(scopes),
),
account=SimpleNamespace(id="account-1"),
user=SimpleNamespace(id="user-1"),
)
if __name__ == "__main__":
unittest.main()
+468
View File
@@ -0,0 +1,468 @@
from __future__ import annotations
import unittest
from unittest.mock import patch
from sqlalchemy import create_engine
from sqlalchemy.orm import sessionmaker
from govoplan_core.auth import ApiPrincipal
from govoplan_core.core.access import PrincipalRef
from govoplan_core.core.change_sequence import ChangeSequenceEntry
from govoplan_core.core.datasources import (
CAPABILITY_DATASOURCE_ORIGINS,
CAPABILITY_POLICY_DATASOURCE_VISIBILITY,
DatasourceAccessError,
DatasourceGovernance,
DatasourceNotFoundError,
DatasourceField,
DatasourceOrigin,
DatasourceOriginReadResult,
DatasourceReadRequest,
DatasourceVisibilityPolicyDecision,
)
from govoplan_core.db.base import Base, utcnow
from govoplan_datasources.backend.db.models import (
DatasourceGovernanceReferenceRecord,
DatasourceMaterializationRecord,
DatasourcePayloadRecord,
DatasourceRecord,
)
from govoplan_datasources.backend.service import (
CATALOGUE_READ_SCOPE,
SqlDatasourceProvider,
)
SCHEMA = [
{"name": "id", "data_type": "integer", "nullable": False},
{"name": "owner_id", "data_type": "string", "nullable": False},
{"name": "secret", "data_type": "string", "nullable": False},
{"name": "internal_note", "data_type": "string", "nullable": True},
]
ROWS = [
{
"id": 1,
"owner_id": "account-1",
"secret": "protected-one",
"internal_note": "hidden-one",
},
{
"id": 2,
"owner_id": "account-2",
"secret": "protected-two",
"internal_note": "hidden-two",
},
]
def principal(
*,
tenant_id: str = "tenant-1",
account_id: str = "account-1",
role_ids: tuple[str, ...] = ("reader",),
auth_method: str = "session",
service_account_id: str | None = None,
) -> ApiPrincipal:
return ApiPrincipal(
principal=PrincipalRef(
account_id=account_id,
membership_id=(None if auth_method == "service_account" else "member-1"),
tenant_id=tenant_id,
scopes=frozenset({CATALOGUE_READ_SCOPE}),
role_ids=frozenset(role_ids),
auth_method=auth_method,
service_account_id=service_account_id,
),
account=object(),
user=object(),
)
class _PolicyProvider:
def __init__(self, decision: DatasourceVisibilityPolicyDecision) -> None:
self.decision = decision
self.requests = []
def decide_datasource_visibility(self, _session, *, request):
self.requests.append(request)
return self.decision
class _Registry:
def __init__(self, policy_provider: _PolicyProvider) -> None:
self.policy_provider = policy_provider
def has_capability(self, name: str) -> bool:
return name == CAPABILITY_POLICY_DATASOURCE_VISIBILITY
def capability(self, name: str) -> object:
if not self.has_capability(name):
raise KeyError(name)
return self.policy_provider
class _OriginProvider:
origin = DatasourceOrigin(
ref="secret:origin",
source_name="live_cases",
name="Live cases",
kind="database",
shape="tabular",
supported_modes=("live",),
provider="test",
schema=tuple(
DatasourceField(
name=str(field["name"]),
data_type=str(field["data_type"]),
nullable=bool(field["nullable"]),
)
for field in SCHEMA
),
fingerprint="live-base-fingerprint",
row_count=2,
)
def list_origins(self, _session, _principal, *, query="", limit=100):
del query
return (self.origin,)[:limit]
def get_origin(self, _session, _principal, *, origin_ref):
return self.origin if origin_ref == self.origin.ref else None
def read_origin(self, _session, _principal, *, request):
rows = ROWS[request.offset : request.offset + request.limit]
return DatasourceOriginReadResult(
origin=self.origin,
rows=tuple(rows),
total_rows=2,
truncated=request.offset + len(rows) < 2,
elapsed_ms=1,
)
class _LiveRegistry:
def __init__(self) -> None:
self.origin_provider = _OriginProvider()
def has_capability(self, name: str) -> bool:
return name == CAPABILITY_DATASOURCE_ORIGINS
def capability(self, name: str) -> object:
if not self.has_capability(name):
raise KeyError(name)
return self.origin_provider
class DatasourceVisibilityTests(unittest.TestCase):
def setUp(self) -> None:
self.engine = create_engine("sqlite:///:memory:")
Base.metadata.create_all(
self.engine,
tables=[
DatasourceRecord.__table__,
DatasourceGovernanceReferenceRecord.__table__,
DatasourcePayloadRecord.__table__,
DatasourceMaterializationRecord.__table__,
ChangeSequenceEntry.__table__,
],
)
self.Session = sessionmaker(bind=self.engine)
self.session = self.Session()
def tearDown(self) -> None:
self.session.close()
Base.metadata.drop_all(
self.engine,
tables=[
DatasourceMaterializationRecord.__table__,
ChangeSequenceEntry.__table__,
DatasourcePayloadRecord.__table__,
DatasourceGovernanceReferenceRecord.__table__,
DatasourceRecord.__table__,
],
)
self.engine.dispose()
def _datasource(
self,
*,
policy: dict[str, object] | None = None,
access_policy_ref: str | None = None,
frozen: bool = False,
snapshot_policy: dict[str, object] | None = None,
) -> DatasourceRecord:
item = DatasourceRecord(
tenant_id="tenant-1",
source_name="governed_cases",
name="Governed cases",
kind="upload",
mode="static",
shape="tabular",
status="active",
provider="private.provider",
provider_ref="secret:origin",
schema_=SCHEMA,
fingerprint="base-fingerprint",
row_count=2,
byte_count=250,
provenance_={"secret_locator": "private://rows"},
metadata_={"credential_ref": "credential:secret"},
access_policy_ref=access_policy_ref,
visibility_policy=policy or {},
)
self.session.add(item)
self.session.flush()
snapshot = DatasourceGovernance(
publication_state="internal",
visibility_policy=snapshot_policy or policy or {},
access_policy_ref=access_policy_ref,
)
materialization = DatasourceMaterializationRecord(
tenant_id="tenant-1",
datasource_id=item.id,
revision=1,
state="published",
schema_=SCHEMA,
rows=ROWS,
fingerprint="materialization-fingerprint",
row_count=2,
byte_count=250,
frozen_at=utcnow() if frozen else None,
frozen_label="Evidence" if frozen else None,
provenance_={"protected": "source details"},
metadata_={"protected": "materialization details"},
governance_snapshot_=snapshot.to_dict(),
)
self.session.add(materialization)
self.session.flush()
item.current_materialization_id = materialization.id
self.session.flush()
return item
def test_role_acl_filters_discovery_and_denies_reads(self) -> None:
item = self._datasource(policy={"source_acl": {"role_ids": ["reader"]}})
provider = SqlDatasourceProvider()
self.assertEqual(1, len(provider.list_datasources(self.session, principal())))
denied = principal(role_ids=("other",))
self.assertEqual((), provider.list_datasources(self.session, denied))
self.assertIsNone(
provider.get_datasource(
self.session,
denied,
datasource_ref=f"datasource:{item.id}",
)
)
with self.assertRaises(DatasourceAccessError):
provider.read_datasource(
self.session,
denied,
request=DatasourceReadRequest(datasource_ref=f"datasource:{item.id}"),
)
other_tenant = principal(tenant_id="tenant-2")
self.assertEqual((), provider.list_datasources(self.session, other_tenant))
with self.assertRaises(DatasourceNotFoundError):
provider.read_datasource(
self.session,
other_tenant,
request=DatasourceReadRequest(datasource_ref=f"datasource:{item.id}"),
)
def test_rows_and_fields_are_filtered_into_an_opaque_permitted_view(self) -> None:
policy = {
"source_acl": {"role_ids": ["reader"]},
"fields": {
"secret": {
"classification": "restricted",
"action": "redact",
"allow": {"role_ids": ["privileged"]},
},
"internal_note": {
"classification": "confidential",
"action": "omit",
"allow": {"role_ids": ["privileged"]},
},
},
"row_filters": [
{"field": "owner_id", "claim": "account_id", "operator": "equals"}
],
}
item = self._datasource(policy=policy)
provider = SqlDatasourceProvider()
with patch(
"govoplan_datasources.backend.service.audit_event"
) as audit_event:
result = provider.read_datasource(
self.session,
principal(),
request=DatasourceReadRequest(
datasource_ref=f"datasource:{item.id}"
),
)
self.assertEqual(
({"id": 1, "owner_id": "account-1", "secret": None},),
result.rows,
)
self.assertEqual(
["id", "owner_id", "secret"],
[field.name for field in result.datasource.schema],
)
self.assertEqual("restricted", result.datasource.schema[-1].classification)
self.assertNotEqual(
"materialization-fingerprint", result.datasource.fingerprint
)
self.assertIsNone(result.datasource.provider)
self.assertIsNone(result.datasource.provider_ref)
self.assertNotIn("credential_ref", result.datasource.metadata)
self.assertEqual(1, result.total_rows)
self.assertEqual(
["datasource.visibility_applied"],
[diagnostic.code for diagnostic in result.diagnostics],
)
audit_details = audit_event.call_args.kwargs["details"]
self.assertEqual(2, audit_details["visibility_summary"]["field_rule_count"])
self.assertNotIn("protected-one", str(audit_details))
repeated = provider.read_datasource(
self.session,
principal(),
request=DatasourceReadRequest(
datasource_ref=f"datasource:{item.id}",
expected_fingerprint=result.datasource.fingerprint,
),
)
self.assertEqual(result.datasource.fingerprint, repeated.datasource.fingerprint)
def test_materialization_acl_supports_service_principals(self) -> None:
item = self._datasource(
policy={
"source_acl": {"auth_methods": ["session", "service_account"]},
"materialization_acl": {"service_account_ids": ["service:reporting"]},
}
)
provider = SqlDatasourceProvider()
with self.assertRaises(DatasourceAccessError):
provider.read_datasource(
self.session,
principal(),
request=DatasourceReadRequest(datasource_ref=f"datasource:{item.id}"),
)
result = provider.read_datasource(
self.session,
principal(
account_id="service-account",
auth_method="service_account",
service_account_id="service:reporting",
),
request=DatasourceReadRequest(datasource_ref=f"datasource:{item.id}"),
)
self.assertEqual(2, result.total_rows)
def test_frozen_state_keeps_its_restrictive_local_policy(self) -> None:
snapshot_policy = {"source_acl": {"role_ids": ["evidence-reader"]}}
item = self._datasource(
policy={},
frozen=True,
snapshot_policy=snapshot_policy,
)
provider = SqlDatasourceProvider()
with self.assertRaises(DatasourceAccessError):
provider.read_datasource(
self.session,
principal(role_ids=("reader",)),
request=DatasourceReadRequest(
datasource_ref=f"datasource:{item.id}",
consistency="frozen",
),
)
result = provider.read_datasource(
self.session,
principal(role_ids=("evidence-reader",)),
request=DatasourceReadRequest(
datasource_ref=f"datasource:{item.id}",
consistency="frozen",
),
)
self.assertEqual(2, result.total_rows)
def test_external_policy_tightens_local_behavior_and_missing_provider_denies(
self,
) -> None:
item = self._datasource(access_policy_ref="case-workers")
with self.assertRaises(DatasourceAccessError):
SqlDatasourceProvider().read_datasource(
self.session,
principal(),
request=DatasourceReadRequest(datasource_ref=f"datasource:{item.id}"),
)
external = _PolicyProvider(
DatasourceVisibilityPolicyDecision(
allowed=True,
policies=({"source_acl": {"role_ids": ["case-worker"]}},),
decision_ref="policy-decision:1",
)
)
provider = SqlDatasourceProvider(registry=_Registry(external))
result = provider.read_datasource(
self.session,
principal(role_ids=("case-worker",)),
request=DatasourceReadRequest(datasource_ref=f"datasource:{item.id}"),
)
self.assertEqual(2, result.total_rows)
self.assertEqual("case-workers", external.requests[-1].policy_ref)
def test_denied_audit_contains_no_protected_values(self) -> None:
item = self._datasource(policy={"source_acl": {"role_ids": ["authorized"]}})
provider = SqlDatasourceProvider()
with patch("govoplan_datasources.backend.service.audit_event") as audit_event:
with self.assertRaises(DatasourceAccessError):
provider.read_datasource(
self.session,
principal(role_ids=("denied",)),
request=DatasourceReadRequest(
datasource_ref=f"datasource:{item.id}"
),
)
details = audit_event.call_args.kwargs["details"]
self.assertEqual("denied", details["outcome"])
serialized = str(details)
self.assertNotIn("protected-one", serialized)
self.assertNotIn("hidden-one", serialized)
self.assertNotIn("credential:secret", serialized)
def test_live_rows_are_filtered_before_origin_data_is_returned(self) -> None:
item = self._datasource(
policy={
"source_acl": {"role_ids": ["reader"]},
"row_filters": [
{"field": "owner_id", "claim": "account_id"}
],
}
)
item.mode = "live"
self.session.flush()
result = SqlDatasourceProvider(registry=_LiveRegistry()).read_datasource(
self.session,
principal(),
request=DatasourceReadRequest(
datasource_ref=f"datasource:{item.id}",
consistency="live",
),
)
self.assertEqual(1, result.total_rows)
self.assertEqual("account-1", result.rows[0]["owner_id"])
self.assertNotEqual("live-base-fingerprint", result.datasource.fingerprint)
self.assertIsNone(result.datasource.provider_ref)
if __name__ == "__main__":
unittest.main()
+1 -1
View File
@@ -1,6 +1,6 @@
{ {
"name": "@govoplan/datasources-webui", "name": "@govoplan/datasources-webui",
"version": "0.1.18", "version": "0.1.21",
"private": true, "private": true,
"type": "module", "type": "module",
"main": "src/index.ts", "main": "src/index.ts",
+114
View File
@@ -28,11 +28,15 @@ export type DatasourceGovernance = {
classification: string; classification: string;
privacy_profile_ref?: string | null; privacy_profile_ref?: string | null;
retention_policy_ref?: string | null; retention_policy_ref?: string | null;
access_policy_ref?: string | null;
visibility_policy: Record<string, unknown>;
hold_refs: string[]; hold_refs: string[];
publication_state: string; publication_state: string;
transfer_agreement_ref?: string | null; transfer_agreement_ref?: string | null;
freshness_policy: Record<string, unknown>; freshness_policy: Record<string, unknown>;
quality_policy: Record<string, unknown>; quality_policy: Record<string, unknown>;
approval_policy: Record<string, unknown>;
retention_policy: Record<string, unknown>;
known_limits: string[]; known_limits: string[];
correction_procedure_ref?: string | null; correction_procedure_ref?: string | null;
affected_refs: string[]; affected_refs: string[];
@@ -43,6 +47,7 @@ export type DatasourceField = {
name: string; name: string;
data_type: string; data_type: string;
nullable: boolean; nullable: boolean;
classification: string;
}; };
export type Datasource = { export type Datasource = {
@@ -82,6 +87,8 @@ export type DatasourceMaterialization = {
frozen_label?: string | null; frozen_label?: string | null;
source_timestamp?: string | null; source_timestamp?: string | null;
created_at?: string | null; created_at?: string | null;
disposed_at?: string | null;
disposition: Record<string, unknown>;
provenance: Record<string, unknown>; provenance: Record<string, unknown>;
metadata: Record<string, unknown>; metadata: Record<string, unknown>;
governance: DatasourceGovernance; governance: DatasourceGovernance;
@@ -139,6 +146,21 @@ export type DatasourceStage = {
row_count?: number | null; row_count?: number | null;
byte_count?: number | null; byte_count?: number | null;
validation: DatasourceStageValidation; validation: DatasourceStageValidation;
approval: {
state?: "not_required" | "pending" | "approved" | "rejected" | "expired";
policy_hash?: string;
subject_digest?: string;
required_approvals?: number;
approval_count?: number;
expires_at?: string | null;
approvals?: Array<{
actor_ref?: string;
decision?: "approve" | "reject";
reason?: string;
decided_at?: string;
}>;
policy?: Record<string, unknown>;
};
created_at?: string | null; created_at?: string | null;
promoted_at?: string | null; promoted_at?: string | null;
promoted_materialization_ref?: string | null; promoted_materialization_ref?: string | null;
@@ -200,6 +222,38 @@ export type DatasourcePreview = {
}>; }>;
}; };
export type DatasourceLifecycleEvidence = {
ref: string;
subject_ref: string;
event_type: string;
occurred_at: string;
actor_ref?: string | null;
policy_version?: string | null;
policy_hash?: string | null;
subject_digest: string;
previous_event_hash?: string | null;
event_hash: string;
details: Record<string, unknown>;
};
export type DatasourceRetentionCandidate = {
ref: string;
kind: "stage" | "materialization";
datasource_ref?: string | null;
disposition: "delete_stage" | "purge_materialization_payload";
eligible_at: string;
eligible: boolean;
blockers: string[];
policy_version: string;
policy_hash: string;
};
export type DatasourceRetentionPlan = {
as_of: string;
plan_hash: string;
candidates: DatasourceRetentionCandidate[];
};
export async function listDatasources( export async function listDatasources(
settings: ApiSettings, settings: ApiSettings,
query = "", query = "",
@@ -290,6 +344,22 @@ export function promoteDatasourceStage(
}); });
} }
export function decideDatasourceStage(
settings: ApiSettings,
stageRef: string,
payload: {
decision: "approve" | "reject";
reason: string;
expected_policy_hash: string;
expected_subject_digest: string;
}
): Promise<DatasourceStage> {
return apiFetch(settings, `/api/v1/datasources/stages/${refId(stageRef)}/decision`, {
method: "POST",
body: JSON.stringify(payload)
});
}
export function registerDatasourceOrigin( export function registerDatasourceOrigin(
settings: ApiSettings, settings: ApiSettings,
payload: { payload: {
@@ -316,6 +386,50 @@ export function refreshDatasource(
}); });
} }
export function prepareDatasourceRefresh(
settings: ApiSettings,
datasourceRef: string
): Promise<DatasourceStage> {
return apiFetch(settings, `/api/v1/datasources/${refId(datasourceRef)}/refresh/stage`, {
method: "POST"
});
}
export async function listDatasourceLifecycleEvidence(
settings: ApiSettings,
subjectRef?: string
): Promise<DatasourceLifecycleEvidence[]> {
const params = new URLSearchParams();
if (subjectRef) params.set("subject_ref", subjectRef);
const suffix = params.size ? `?${params.toString()}` : "";
const response = await apiFetch<{ evidence: DatasourceLifecycleEvidence[] }>(
settings,
`/api/v1/datasources/lifecycle-evidence${suffix}`
);
return response.evidence;
}
export function previewDatasourceRetention(
settings: ApiSettings
): Promise<DatasourceRetentionPlan> {
return apiFetch(settings, "/api/v1/datasources/retention/plan");
}
export function applyDatasourceRetention(
settings: ApiSettings,
plan: DatasourceRetentionPlan,
targetRefs: string[]
): Promise<{ plan_hash: string; disposed_refs: string[]; evidence_hashes: string[] }> {
return apiFetch(settings, "/api/v1/datasources/retention/apply", {
method: "POST",
body: JSON.stringify({
as_of: plan.as_of,
plan_hash: plan.plan_hash,
target_refs: targetRefs
})
});
}
export function freezeDatasource( export function freezeDatasource(
settings: ApiSettings, settings: ApiSettings,
datasourceRef: string, datasourceRef: string,
File diff suppressed because it is too large Load Diff
@@ -15,6 +15,11 @@ export const DATASOURCE_GOVERNANCE_DOCUMENTATION = {
documentationType: "admin" documentationType: "admin"
} satisfies DocumentationHelpReference; } satisfies DocumentationHelpReference;
export const DATASOURCE_VISIBILITY_DOCUMENTATION = {
topicId: "datasources.visibility",
documentationType: "admin"
} satisfies DocumentationHelpReference;
export const DATASOURCES_I18N = { export const DATASOURCES_I18N = {
loading: "i18n:govoplan-datasources.loading_reason", loading: "i18n:govoplan-datasources.loading_reason",
working: "i18n:govoplan-datasources.working_reason", working: "i18n:govoplan-datasources.working_reason",
+8
View File
@@ -63,6 +63,10 @@ const en = {
"Semantic definition": "Semantic definition", "Semantic definition": "Semantic definition",
"Freshness policy (JSON)": "Freshness policy (JSON)", "Freshness policy (JSON)": "Freshness policy (JSON)",
"Quality policy (JSON)": "Quality policy (JSON)", "Quality policy (JSON)": "Quality policy (JSON)",
"Access policy reference": "Access policy reference",
"Optional Policy module target": "Optional Policy module target",
"Visibility policy (JSON)": "Visibility policy (JSON)",
"Configure source and materialization ACLs, field redaction or omission, and principal-bound row filters.": "Configure source and materialization ACLs, field redaction or omission, and principal-bound row filters.",
"Quality policy": "Quality policy", "Quality policy": "Quality policy",
"Policy hash": "Policy hash", "Policy hash": "Policy hash",
"Schema change": "Schema change", "Schema change": "Schema change",
@@ -136,6 +140,10 @@ const de: Record<keyof typeof en, string> = {
"Semantic definition": "Semantische Definition", "Semantic definition": "Semantische Definition",
"Freshness policy (JSON)": "Aktualitätsrichtlinie (JSON)", "Freshness policy (JSON)": "Aktualitätsrichtlinie (JSON)",
"Quality policy (JSON)": "Qualitätsrichtlinie (JSON)", "Quality policy (JSON)": "Qualitätsrichtlinie (JSON)",
"Access policy reference": "Zugriffsrichtlinienreferenz",
"Optional Policy module target": "Optionales Ziel im Richtlinienmodul",
"Visibility policy (JSON)": "Sichtbarkeitsrichtlinie (JSON)",
"Configure source and materialization ACLs, field redaction or omission, and principal-bound row filters.": "Konfigurieren Sie Zugriffslisten für Quellen und Materialisierungen, Feldschwärzung oder -auslassung sowie an den Akteur gebundene Zeilenfilter.",
"Quality policy": "Qualitätsrichtlinie", "Quality policy": "Qualitätsrichtlinie",
"Policy hash": "Richtlinien-Hash", "Policy hash": "Richtlinien-Hash",
"Schema change": "Schemaänderung", "Schema change": "Schemaänderung",
+3 -2
View File
@@ -10,14 +10,15 @@ const readScopes = ["datasources:catalogue:read", "datasources:source:admin"];
export const datasourcesModule: PlatformWebModule = { export const datasourcesModule: PlatformWebModule = {
id: "datasources", id: "datasources",
label: "i18n:govoplan-datasources.datasources", label: "i18n:govoplan-datasources.datasources",
version: "0.1.14", version: "0.1.18",
optionalDependencies: [ optionalDependencies: [
"access", "access",
"audit", "audit",
"connectors", "connectors",
"files", "files",
"notifications", "notifications",
"policy" "policy",
"search"
], ],
translations: generatedTranslations, translations: generatedTranslations,
viewSurfaces: [ viewSurfaces: [
+13 -209
View File
@@ -1,65 +1,9 @@
.datasources-page {
position: relative;
height: calc(100vh - 115px);
min-width: 0;
min-height: 0;
padding: 0;
overflow: hidden;
color: var(--text);
background: var(--bg);
}
.datasources-page *,
.datasources-page *::before,
.datasources-page *::after {
box-sizing: border-box;
}
.datasources-shell {
display: grid;
grid-template-columns: minmax(250px, 300px) minmax(0, 1fr);
width: 100%;
height: 100%;
min-width: 0;
min-height: 0;
overflow: hidden;
border: var(--border-line);
background: var(--panel);
}
.datasources-sidebar,
.datasources-workspace,
.datasources-content, .datasources-content,
.datasources-list-frame { .datasources-list-frame {
min-width: 0; min-width: 0;
min-height: 0; min-height: 0;
} }
.datasources-sidebar {
display: flex;
flex-direction: column;
overflow: hidden;
border-right: var(--border-line);
background: var(--panel-soft);
}
.datasources-sidebar-toolbar,
.datasources-workspace-toolbar,
.datasources-section-heading {
display: flex;
align-items: center;
justify-content: space-between;
gap: 10px;
flex: 0 0 auto;
border-bottom: var(--border-line);
background: var(--panel-header);
}
.datasources-sidebar-toolbar {
min-height: 52px;
padding: 8px 10px 8px 14px;
}
.datasources-toolbar-actions, .datasources-toolbar-actions,
.datasources-current-title, .datasources-current-title,
.datasources-current-title > span, .datasources-current-title > span,
@@ -75,80 +19,11 @@
background: var(--panel); background: var(--panel);
} }
.datasources-search {
padding: 9px;
border-bottom: var(--border-line);
}
.datasources-search input {
width: 100%;
min-height: 34px;
padding: 7px 9px;
}
.datasources-list-frame { .datasources-list-frame {
flex: 1 1 auto; flex: 1 1 auto;
overflow: hidden; overflow: hidden;
} }
.datasources-list {
height: 100%;
overflow: auto;
padding: 6px;
}
.datasources-list > button {
display: grid;
grid-template-columns: minmax(0, 1fr) auto;
align-items: center;
gap: 8px;
width: 100%;
min-height: 56px;
padding: 8px 9px;
border: 0;
border-radius: var(--radius-sm);
background: transparent;
color: var(--text);
cursor: pointer;
text-align: left;
}
.datasources-list > button:hover,
.datasources-list > button:focus-visible {
background: var(--primary-soft);
outline: none;
}
.datasources-list > button.is-selected {
background: var(--primary-soft-strong);
box-shadow: inset 3px 0 0 var(--accent);
}
.datasources-list > button > span:first-child {
min-width: 0;
}
.datasources-list strong,
.datasources-list small {
display: block;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.datasources-list strong {
color: var(--text-strong);
font-size: 13px;
}
.datasources-list small {
margin-top: 4px;
color: var(--muted);
font-size: 11px;
}
.datasources-list-empty,
.datasources-inline-empty,
.datasources-inline-loading { .datasources-inline-loading {
display: grid; display: grid;
place-items: center; place-items: center;
@@ -158,17 +33,8 @@
text-align: center; text-align: center;
} }
.datasources-workspace {
position: relative;
display: flex;
flex-direction: column;
overflow: hidden;
background: var(--bg);
}
.datasources-workspace-toolbar { .datasources-workspace-toolbar {
min-height: 58px; min-width: 0;
padding: 8px 10px 8px 14px;
} }
.datasources-current-title { .datasources-current-title {
@@ -218,43 +84,6 @@
padding: 14px; padding: 14px;
} }
.datasources-metrics {
display: grid;
grid-template-columns: repeat(5, minmax(100px, 1fr));
gap: 8px;
margin-bottom: 14px;
}
.datasources-metrics > span {
min-width: 0;
min-height: 61px;
padding: 9px 10px;
border: var(--border-line);
border-radius: var(--radius-sm);
background: var(--panel);
}
.datasources-metrics small,
.datasources-metrics strong {
display: block;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.datasources-metrics small {
color: var(--muted);
font-size: 10px;
text-transform: uppercase;
}
.datasources-metrics strong {
margin-top: 6px;
color: var(--text-strong);
font-size: 14px;
text-transform: capitalize;
}
.datasources-description { .datasources-description {
margin-bottom: 14px; margin-bottom: 14px;
padding: 10px 12px; padding: 10px 12px;
@@ -265,18 +94,6 @@
line-height: 1.45; line-height: 1.45;
} }
.datasources-detail-section {
min-width: 0;
margin-bottom: 14px;
border: var(--border-line);
background: var(--panel);
}
.datasources-section-heading {
min-height: 42px;
padding: 7px 10px;
}
.datasources-section-heading > span { .datasources-section-heading > span {
color: var(--text-strong); color: var(--text-strong);
font-size: 13px; font-size: 13px;
@@ -470,6 +287,17 @@
max-height: min(860px, calc(100vh - 32px)); max-height: min(860px, calc(100vh - 32px));
} }
.datasources-retention-select-cell {
width: 2.5rem;
text-align: center;
}
.datasources-table-scroll td small {
display: block;
margin-top: 0.2rem;
color: var(--color-text-muted);
}
.datasources-governance-dialog .datasources-dialog-fields textarea { .datasources-governance-dialog .datasources-dialog-fields textarea {
min-height: 76px; min-height: 76px;
} }
@@ -487,12 +315,6 @@
font-size: 12px; font-size: 12px;
} }
.datasources-dialog-grid {
display: grid;
grid-template-columns: minmax(0, 1fr) minmax(0, 1fr);
gap: 12px;
}
.datasources-dialog-copy { .datasources-dialog-copy {
margin: 0 0 14px; margin: 0 0 14px;
color: var(--muted); color: var(--muted);
@@ -507,26 +329,10 @@
overflow: auto; overflow: auto;
} }
.datasources-shell {
grid-template-columns: 1fr;
height: auto;
overflow: visible;
}
.datasources-sidebar {
max-height: 42vh;
border-right: 0;
border-bottom: var(--border-line);
}
.datasources-workspace { .datasources-workspace {
min-height: 58vh; min-height: 58vh;
} }
.datasources-metrics {
grid-template-columns: repeat(2, minmax(0, 1fr));
}
.datasources-workspace-toolbar { .datasources-workspace-toolbar {
align-items: flex-start; align-items: flex-start;
flex-wrap: wrap; flex-wrap: wrap;
@@ -538,9 +344,7 @@
} }
@media (max-width: 560px) { @media (max-width: 560px) {
.datasources-metrics, .datasources-key-values {
.datasources-key-values,
.datasources-dialog-grid {
grid-template-columns: 1fr; grid-template-columns: 1fr;
} }