On this page
Research question
Can attention retrieval and residual control be separated so that a learned sinusoidal control path determines how retrieved information influences the residual stream?
Alethic is my language-model research line. The current direction is Sin-Gated Control Attention (SGCA). The question concerns optimization, representation, and efficiency under a carefully matched decoder-only baseline.
Architecture
The decoder block uses RMSNorm, GQA, SGCA, and SwiGLU. RoPE remains in the attention path.
is the output of GQA. The learned tensor must be broadcast-compatible with ; no scalar or fixed matrix shape is assumed. The block then applies RMSNorm, SwiGLU, and another residual update.
Matched baseline
Both paths share GQA. The baseline adds the attention output directly; SGCA computes a separate control update.
# Matched GQA baseline
H = rms_norm(X)
A = gqa(H)
X_prime = X + A
output = X_prime + swiglu(rms_norm(X_prime))
# SGCA research block
H = rms_norm(X)
A = gqa(H)
C = linear(H) * (alpha * sin(A))
X_prime = X + C
output = X_prime + swiglu(rms_norm(X_prime))This is conceptual pseudocode, not a training implementation. The actual tensor shape and initialization of α belong in each run configuration.
Experiment methodology
Match the tokenizer, corpus, data order, optimizer, learning-rate schedule, token budget, context, precision, batch construction, initialization, hardware class, and evaluation where possible. Parameter matching and compute matching answer different questions and must be reported explicitly. Multiple seeds and non-sinusoidal gates are part of the validation plan.
Model lineage
- SinGatedLMPreliminary ~64K and ~1M controlled experiments.
- 116M pretrainingHistorical predecessor released under the Ruzz name.
- Alethic-151MExperimental pretraining and SinGatedAttention exploration.
- SGCACurrent separate sinusoidal residual-control formulation.
- Phase IPlanned ~1B scale, 8,240-token context, controlled baseline.
This is a research progression, not one benchmark series. The models did not all share the same architecture or setup.
Recorded evidence
The earlier parameter-matched results are documented on the SinGatedLM page. A later Alethic-151M record reports a best validation loss of approximately 5.0279. These separate experiments are not an apples-to-apples comparison with current SGCA. No Phase-I SGCA result is reported here.
Current experiment
| Configuration | Phase-I target |
|---|---|
| Parameters | Approximately 1B, planned |
| Context | 8,240 tokens |
| Tokenizer | Approximately 48K vocabulary, planned |
| Training tokens | Approximately 20B for a full run, planned |
| Attention / position | GQA / RoPE in attention path |
| Control | SGCA with learned broadcast-compatible α |
| Normalization / FFN | RMSNorm / SwiGLU |
Earlier 1.5B and 12K/54K plans are historical. Longer-context scaling is postponed until the smaller-scale mechanism has been evaluated.
Compute and pipeline
Work has used free Colab T4 sessions and Kaggle environments, including 2×T4. Tokenizer preparation, export, source-ID deduplication, and binary packing are substantial parts of the infrastructure. The challenge is making runs resumable and comparable under constrained compute.
Limitations
Small-model evidence does not establish SGCA superiority. Earlier experiments have limited seeds, dataset breadth, and model scale. Training stability, initialization sensitivity, and throughput remain questions to measure.
Reproducibility
The earlier code and model releases are linked below. A public Phase-I repository, complete run configuration, seeds, tokenizer release, and final logs have not been supplied for this page.
Updates
1 October 2026: Overview documented with the current Phase-I target and separate historical model lineage.