security-ir: experimental 384D checkpoint
Held-out exact IR reconstruction: 0/2. Native structural validity: 2/2. These two authored paraphrase examples share semantic targets with training; they are not independent real-world benchmark tasks. The model generates plausible but semantically incorrect candidates and must not authorize actions.
This is an opt-in development artifact. It is not the default symbolic compiler, not a verified formalizer, and not a security repair policy. All qualification, admission and proof-authority flags remain false. No LLM/provider fallback runs.
Runtime
Input embeddings use thenlper/gte-small at commit
17e1f347d17fe144873b1201da91788898c639cd (384 dimensions). The embedding encoder
weights are not included. Source inference requires the corresponding verified
local encoder assets and rejects inputs above 512 tokens without truncation.
The installed ipfs_datasets_py implementation must match the source hashes in
the checkpoint; incompatible sources are rejected. This development runtime is
currently implemented in the matching workspace checkout, not claimed to be in
an existing PyPI release. No repository-supplied Python is executed on download.
from ipfs_datasets_py.logic.security_ir import open_autoencoder
model = open_autoencoder() # library catalog pins a complete Hub commit + SHA256
report = model.infer([{"id": "sample", "source_text": source, "embedding": vector384}])
# Or: model.infer_texts([source], snapshot_path=verified_gte_snapshot)
Training and scope
Security, Intent and UI/UX heads inherit all compatible nonlexical tensors from a trained local Legal 384D head. New lexical rows use trained parent row means; no random parameters remain in their initial transferred state. Each native head used 12 authored training rows, 2 tuning rows, 500 epochs / 1000 optimizer updates; selection used tuning loss only. Outputs are explicitly local typed fragments, not complete verified programs or native documents. Security operator fragments still require their enclosing program and source-reference context. Legal retains its separate sparse core and learned formula head, with parser-derived features.
These checkpoints were not trained on the complete CVEFixes or SkillCenter corpora. The separate full-feature structural-family experiment (15/9/8/6 routes) does not measure this 384D source decoder and is not its training loss or coverage. Weight/conditioning ablations demonstrate numerical dependence, not correct meaning.
Provenance and license
The manifest records the actual parent hashes, training manifests, embedding assets and measured validation. Source screening found no private documents, credentials, benchmark solutions or third-party corpus text in the new authored fixture export. Original Legal files and the upstream Hugging Face datasets were not overwritten. This release makes no new license grant for inherited weights. The runtime source repository retains its AGPL-3.0 license; the new authored fixture export separately declares Apache-2.0. A corpus license is not asserted to be the weight license.
The checkpoint and evidence are immutable under releases/20261002-384-development-v1/.
Manifest SHA256: d9fc241c7c6a039068681fab18b6a36fbcbd1752c2683ba1dc16423ce94aba2a.
Additional verified development releases (2026-10-01)
The following opt-in structured profiles use a fixed typed JSON schema and learned scalar vocabulary. Their cards distinguish raw reconstruction, hybrid AST normalization, and compilation. They do not replace the symbolic formalizers or confer proof authority.
Runtime code and exact release descriptors. All new package hashes and cold/offline numerical replay were verified.