An Attempt at Bridging Human Cognition and AI: Integrating Titan with Transformer²

| January 18, 2025

A design sketch for combining Titan's surprise-based memory with Transformer² SVD fine-tuning, with the math worked out and no experiments run yet.

This post contains affiliate links for tools I use in production. If you buy through them I earn a commission at no extra cost to you. Recommendations are based on my own experience.

Table of Contents

Why integrate Titan and Transformer²?

Think about how you recall a meeting: you forget most of the back-and-forth but keep the decisions. Human memory compresses aggressively, keeping the salient parts and discarding the rest. That is the property this post tries to borrow for a memory module design.

The idea below is a design sketch, not a validated result. I have not run the experiments (no stable GPU access at the time of writing), so treat the following as a hypothesis with the math worked out, not a benchmarked claim.

The plan is to use Singular Value Decomposition (SVD) inside a Titan-style memory module so the model keeps the dominant patterns in its parameters and discards the rest, the same tradeoff a compressed memory makes.

Using SVD in a Titan-style memory module should, in principle:

  • Compress information: SVD keeps the dominant singular directions of a weight update and drops the rest, similar to how a compressed memory keeps only the salient trace.
  • Improve efficiency: low-rank updates are cheaper to store and apply than full-rank ones.

Review

Titan’s surprise-driven memory update

Titan (“Titans: Learning to Memorize at Test Time”, linked under Related Work below) updates a long-term memory module based on “surprise”: how unexpected new information is relative to what the memory already predicts. For an input token xtx_t, Titan projects it into key-value pairs:

kt=xtWK,vt=xtWV,k_t = x_tW_K, \quad v_t = x_tW_V,

and defines an associative memory loss:

ℓ(Mt−1;xt)=∥Mt−1(kt)−vt∥22.\ell(M_{t-1}; x_t) = \big\|M_{t-1}(k_t) - v_t\big\|_2^2.

Using this loss, Titan updates a “surprise” momentum StS_t with forgetting:

St=ηtSt−1−θt∇ℓ(Mt−1;xt),S_t = \eta_t S_{t-1} - \theta_t \nabla \ell(M_{t-1}; x_t),

and then updates the memory with a forgetting gate:

Mt=(1−αt)Mt−1+St.M_t = (1 - \alpha_t)M_{t-1} + S_t.

This lets Titan encode surprising information while a forgetting gate lets ordinary, expected inputs decay from memory.


Transformer² Singular Value Fine-tuning

Transformer² (“Transformer²: Self-Adaptive LLMs”, also linked below) fine-tunes a weight matrix by scaling its singular values rather than updating the full matrix. For a weight matrix W∈Rd×dW \in \mathbb{R}^{d \times d}, SVD gives:

W=USVT,W = U S V^{T},

where

  • U∈Rd×rU \in \mathbb{R}^{d \times r},
  • S∈Rr×rS \in \mathbb{R}^{r \times r},
  • V∈Rd×rV \in \mathbb{R}^{d \times r},

with rr the rank of the decomposition.

This low-rank factorization cuts the number of trainable parameters while keeping the dominant structure of the weight matrix. Applied to a Titan-style memory module, the memory parameters become the set

θt={Ut,St,Vt},\theta_t = \{U_t, S_t, V_t\},

which evolve over time as the model trains.


A design for integrating Titan and Transformer²

Combining Titan’s memory update with Transformer²’s low-rank parameterization gives a framework where an RL policy adjusts low-rank memory factors based on reward. Here is how that combination works out mathematically.

State, action, and transitions

At time tt, the state is:

st=(xt, Mt, θt),s_t = \Big(x_t,\, M_t,\, \theta_t\Big),

where θt={Ut,St,Vt}\theta_t = \{U_t, S_t, V_t\} are the decomposed memory parameters.

The action ata_t adjusts these factors:

at=(ΔUt, ΔSt, ΔVt).a_t = \big(\Delta U_t,\, \Delta S_t,\, \Delta V_t\big).

The transition updates the parameters:

Ut+1=Ut+ΔUt,St+1=St+ΔSt,Vt+1=Vt+ΔVt,\begin{aligned} U_{t+1} &= U_t + \Delta U_t, \\ S_{t+1} &= S_t + \Delta S_t, \\ V_{t+1} &= V_t + \Delta V_t, \end{aligned}

which implies

Wt+1=Ut+1St+1Vt+1T.W_{t+1} = U_{t+1}S_{t+1}V_{t+1}^{T}.

The memory itself still updates via the surprise mechanism, applied with the next step’s decomposed parameters:

St+1=ηt+1St−θt+1∇ℓ(Mt;xt+1),Mt+1=(1−αt+1)Mt+St+1.\begin{aligned} S_{t+1} &= \eta_{t+1} S_t - \theta_{t+1} \nabla \ell(M_t; x_{t+1}), \\ M_{t+1} &= (1-\alpha_{t+1})M_t + S_{t+1}. \end{aligned}

Policy and reward

A policy πϕ(a∣s)\pi_{\phi}(a \mid s) decides how to adjust the low-rank factors given the current state. The objective is to maximize expected cumulative reward:

J(ϕ)=Eτ∼πϕ[∑t=0Tγtrt],J(\phi) = \mathbb{E}_{\tau \sim \pi_{\phi}}\Bigg[\sum_{t=0}^{T} \gamma^t r_t\Bigg],

where rtr_t is the reward at time tt, reflecting how well the model’s memory helps it perform on the task at hand.

Policy gradient update

By the policy gradient theorem, the update for policy parameters ϕ\phi is:

∇ϕJ(ϕ)=Eπϕ[∑t=0T∇ϕlog⁡πϕ(at∣st)⋅Gt],\nabla_{\phi} J(\phi) = \mathbb{E}_{\pi_{\phi}}\Bigg[\sum_{t=0}^{T} \nabla_{\phi} \log \pi_{\phi}(a_t|s_t) \cdot G_t\Bigg],

with

Gt=∑t′=tTγ t′−trt′.G_t = \sum_{t'=t}^{T} \gamma^{\,t'-t}r_{t'}.

The policy update rule becomes:

ϕ←ϕ+α ∇ϕlog⁡πϕ(at∣st)⋅Gt.\phi \leftarrow \phi + \alpha \, \nabla_{\phi} \log \pi_{\phi}(a_t|s_t) \cdot G_t.

Actions sampled from this policy update the SVD factors, which is the mechanism that ties memory adjustments to long-term reward instead of a fixed learning rule.


Conclusion

Combining Titan’s surprise-based memory update with Transformer²’s low-rank SVD parameterization gives a model where memory adjustments are learned via RL rather than fixed by a hand-tuned decay schedule. The math above is self-consistent, but it is a design, not a result: nothing here has been trained or benchmarked.

What would validate this

The next step is running it: fine-tune a small model with this policy, and check whether the learned (ΔU,ΔS,ΔV)(\Delta U, \Delta S, \Delta V) updates actually improve task performance over the fixed forgetting-gate baseline. That experiment has not been run yet due to lack of stable GPU access; if it happens, the results belong in a follow-up post.

Related work

Paper:

Code: SakanaAI’s Transformer² GitHub Repository

Blog: SakanaAI’s Transformer² Blog Post

If you want to run an experiment like this yourself, the tooling below covers the data and observability layer for tracking training runs, and the newsletter is where the results of this one will ship if the GPU access comes through.

Next steps: scaling to production

If you take this into production, these are the pieces I would add first.

  • Supabase Supabase is a hosted Postgres platform with authentication and storage built in. Postgres with pgvector for embeddings, so you do not run a separate vector store.
  • Datadog Datadog aggregates metrics, logs, and traces for infrastructure monitoring. Traces and cost metrics across model calls, so latency and spend are visible per request.
  • Vercel Vercel hosts frontend applications with a global edge network and CI/CD. Deploys the frontend and edge functions that sit in front of the model API.

Deploying generative AI models to production

Get the free playbook on shipping generative AI models to production.