On this page
Abstract
SGCA is an experimental decoder-only language-model mechanism. It retains grouped-query attention as retrieval, then uses its output inside a sinusoidal control term before the residual update. This is a research hypothesis, not a claim of general superiority.
Motivation
The matched baseline adds retrieved information directly to the residual stream. SGCA computes an update from a learned projection of normalized input, modulated by attention output. Retrieval and the rule governing its influence become conceptually distinct.
Definition and mathematical formulation
- : input residual stream.
- : RMS-normalized input.
- : output of GQA.
- : learned linear control projection.
- : elementwise multiplication.
- : a learnable tensor, broadcast-compatible with .
The definition does not fix α to a scalar or matrix. Exact parameterization must be recorded in an implementation.
Block architecture
The block then proceeds through RMSNorm, SwiGLU, and the feed-forward residual.
Baseline
The two paths share a GQA backbone. The control projection changes parameterization and computation; those differences must be accounted for.
Hypotheses
A separate periodic control path may change representations, gradients, or optimization. It could also introduce instability or initialization sensitivity. Both possibilities require experiments.
Experimental controls
Hold the tokenizer, data, order, batching, optimizer, schedule, token budget, context, precision, hardware class, and evaluation fixed where possible. Report parameter and compute differences and repeat across seeds.
Ablation plan
| Comparison | Question | Status |
|---|---|---|
| SGCA / matched GQA | Effect of the complete control path | Planned validation |
| sin / identity | Contribution of the sinusoid | Planned |
| sin / tanh / sigmoid | Periodicity versus other gates | Planned |
| learned / fixed α | Effect of learned modulation | Planned |
| with / without projection | Contribution of WH + b | Planned |
| parameter / geometry / compute matching | Sensitivity to matching method | Planned |
No outcomes are fabricated for these planned comparisons.
Results
No current SGCA benchmark is published here. Earlier SinGatedLM experiments are related motivation and use a predecessor mechanism.
Limitations and open questions
How should α be initialized? How sensitive is the path to attention-output scale? Does the gate help beyond a small dataset? What is its throughput cost at matched compute? Do gains survive multiple seeds and broader evaluation?
Reproduction
Complete public SGCA run artifacts are still needed. The Alethic overview identifies the available predecessor code and checkpoints.
Citation
Research description by Bekhruz Suleyman, updated 1 October 2026. No DOI or peer-reviewed publication is claimed. Cite the page URL and access date when referring to the formulation.