Research record

SGCA

Sin-Gated Control Attention: separating attention retrieval from learned residual control.

On this page

Abstract

SGCA is an experimental decoder-only language-model mechanism. It retains grouped-query attention as retrieval, then uses its output inside a sinusoidal control term before the residual update. This is a research hypothesis, not a claim of general superiority.

Motivation

The matched baseline adds retrieved information directly to the residual stream. SGCA computes an update from a learned projection of normalized input, modulated by attention output. Retrieval and the rule governing its influence become conceptually distinct.

Definition and mathematical formulation

H=RMSNorm(X)H=\mathrm{RMSNorm}(X) A=GQA(H)A=\mathrm{GQA}(H) C=(WH+b)⊙[α⊙sin⁡(A)]C=(WH+b)\odot[\alpha\odot\sin(A)] X′=X+CX'=X+C

  • XX: input residual stream.
  • HH: RMS-normalized input.
  • AA: output of GQA.
  • WH+bWH+b: learned linear control projection.
  • ⊙\odot: elementwise multiplication.
  • α\alpha: a learnable tensor, broadcast-compatible with AA.

The definition does not fix α to a scalar or matrix. Exact parameterization must be recorded in an implementation.

Block architecture

SGCA block architectureX is normalized into H. H branches into GQA producing A and a linear projection WH plus b. The projection multiplies alpha times sine of A to form C. A residual adds X and C to make X prime. RMSNorm and SwiGLU with another residual form output.Input XH = RMSNorm(X)A = GQA(H)Linear: WH + bC = (WH + b) ⊙ [α ⊙ sin(A)]X′ = X + COutput = X′ + SwiGLU(RMSNorm(X′))Residual X
GQA retrieves; the learned tensor α and linear projection control the residual update. RoPE stays in the attention path.

The block then proceeds through RMSNorm, SwiGLU, and the feed-forward residual.

Baseline

Xbase′=X+AX'_{\mathrm{base}}=X+A X′′=X′+SwiGLU(RMSNorm(X′))X''=X'+\mathrm{SwiGLU}(\mathrm{RMSNorm}(X'))

The two paths share a GQA backbone. The control projection changes parameterization and computation; those differences must be accounted for.

Hypotheses

A separate periodic control path may change representations, gradients, or optimization. It could also introduce instability or initialization sensitivity. Both possibilities require experiments.

Experimental controls

Hold the tokenizer, data, order, batching, optimizer, schedule, token budget, context, precision, hardware class, and evaluation fixed where possible. Report parameter and compute differences and repeat across seeds.

Ablation plan

ComparisonQuestionStatus
SGCA / matched GQAEffect of the complete control pathPlanned validation
sin / identityContribution of the sinusoidPlanned
sin / tanh / sigmoidPeriodicity versus other gatesPlanned
learned / fixed αEffect of learned modulationPlanned
with / without projectionContribution of WH + bPlanned
parameter / geometry / compute matchingSensitivity to matching methodPlanned

No outcomes are fabricated for these planned comparisons.

Results

No current SGCA benchmark is published here. Earlier SinGatedLM experiments are related motivation and use a predecessor mechanism.

Limitations and open questions

How should α be initialized? How sensitive is the path to attention-output scale? Does the gate help beyond a small dataset? What is its throughput cost at matched compute? Do gains survive multiple seeds and broader evaluation?

Reproduction

Complete public SGCA run artifacts are still needed. The Alethic overview identifies the available predecessor code and checkpoints.

Citation

Research description by Bekhruz Suleyman, updated 1 October 2026. No DOI or peer-reviewed publication is claimed. Cite the page URL and access date when referring to the formulation.

Last updated 01 Oct 2026Discuss this work