Research record

Alethic

Language-model architecture research, from controlled small-model experiments to a planned ~1B Phase I.

Hugging Face
On this page

Research question

Can attention retrieval and residual control be separated so that a learned sinusoidal control path determines how retrieved information influences the residual stream?

Alethic is my language-model research line. The current direction is Sin-Gated Control Attention (SGCA). The question concerns optimization, representation, and efficiency under a carefully matched decoder-only baseline.

Architecture

The decoder block uses RMSNorm, GQA, SGCA, and SwiGLU. RoPE remains in the attention path.

SGCA block architectureX is normalized into H. H branches into GQA producing A and a linear projection WH plus b. The projection multiplies alpha times sine of A to form C. A residual adds X and C to make X prime. RMSNorm and SwiGLU with another residual form output.Input XH = RMSNorm(X)A = GQA(H)Linear: WH + bC = (WH + b) ⊙ [α ⊙ sin(A)]X′ = X + COutput = X′ + SwiGLU(RMSNorm(X′))Residual X
GQA retrieves; the learned tensor α and linear projection control the residual update. RoPE stays in the attention path.

H=RMSNorm(X),A=GQA(H)H=\mathrm{RMSNorm}(X),\qquad A=\mathrm{GQA}(H) C=(WH+b)⊙[α⊙sin⁡(A)],X′=X+CC=(WH+b)\odot[\alpha\odot\sin(A)],\qquad X'=X+C

AA is the output of GQA. The learned tensor α\alpha must be broadcast-compatible with AA; no scalar or fixed matrix shape is assumed. The block then applies RMSNorm, SwiGLU, and another residual update.

Matched baseline

Both paths share GQA. The baseline adds the attention output directly; SGCA computes a separate control update.

Conceptual block comparison
# Matched GQA baseline
H = rms_norm(X)
A = gqa(H)
X_prime = X + A
output = X_prime + swiglu(rms_norm(X_prime))

# SGCA research block
H = rms_norm(X)
A = gqa(H)
C = linear(H) * (alpha * sin(A))
X_prime = X + C
output = X_prime + swiglu(rms_norm(X_prime))

This is conceptual pseudocode, not a training implementation. The actual tensor shape and initialization of α belong in each run configuration.

Experiment methodology

Match the tokenizer, corpus, data order, optimizer, learning-rate schedule, token budget, context, precision, batch construction, initialization, hardware class, and evaluation where possible. Parameter matching and compute matching answer different questions and must be reported explicitly. Multiple seeds and non-sinusoidal gates are part of the validation plan.

Model lineage

  1. SinGatedLMPreliminary ~64K and ~1M controlled experiments.
  2. 116M pretrainingHistorical predecessor released under the Ruzz name.
  3. Alethic-151MExperimental pretraining and SinGatedAttention exploration.
  4. SGCACurrent separate sinusoidal residual-control formulation.
  5. Phase IPlanned ~1B scale, 8,240-token context, controlled baseline.

This is a research progression, not one benchmark series. The models did not all share the same architecture or setup.

Recorded evidence

The earlier parameter-matched results are documented on the SinGatedLM page. A later Alethic-151M record reports a best validation loss of approximately 5.0279. These separate experiments are not an apples-to-apples comparison with current SGCA. No Phase-I SGCA result is reported here.

Current experiment

ConfigurationPhase-I target
ParametersApproximately 1B, planned
Context8,240 tokens
TokenizerApproximately 48K vocabulary, planned
Training tokensApproximately 20B for a full run, planned
Attention / positionGQA / RoPE in attention path
ControlSGCA with learned broadcast-compatible α
Normalization / FFNRMSNorm / SwiGLU

Earlier 1.5B and 12K/54K plans are historical. Longer-context scaling is postponed until the smaller-scale mechanism has been evaluated.

Compute and pipeline

Work has used free Colab T4 sessions and Kaggle environments, including 2×T4. Tokenizer preparation, export, source-ID deduplication, and binary packing are substantial parts of the infrastructure. The challenge is making runs resumable and comparable under constrained compute.

Limitations

Small-model evidence does not establish SGCA superiority. Earlier experiments have limited seeds, dataset breadth, and model scale. Training stability, initialization sensitivity, and throughput remain questions to measure.

Reproducibility

The earlier code and model releases are linked below. A public Phase-I repository, complete run configuration, seeds, tokenizer release, and final logs have not been supplied for this page.

Updates

1 October 2026: Overview documented with the current Phase-I target and separate historical model lineage.

Last updated 01 Oct 2026Discuss this work