An Attempt at Bridging Human Cognition and AI: Integrating Titan with Transformer²
A design sketch for combining Titan's surprise-based memory with Transformer² SVD fine-tuning, with the math worked out and no experiments run yet.
This post contains affiliate links for tools I use in production. If you buy through them I earn a commission at no extra cost to you. Recommendations are based on my own experience.
Table of Contents
- Why integrate Titan and Transformer²?
- Review
- A design for integrating Titan and Transformer²
- Conclusion
- What would validate this
- Related work
Why integrate Titan and Transformer²?
Think about how you recall a meeting: you forget most of the back-and-forth but keep the decisions. Human memory compresses aggressively, keeping the salient parts and discarding the rest. That is the property this post tries to borrow for a memory module design.
The idea below is a design sketch, not a validated result. I have not run the experiments (no stable GPU access at the time of writing), so treat the following as a hypothesis with the math worked out, not a benchmarked claim.
The plan is to use Singular Value Decomposition (SVD) inside a Titan-style memory module so the model keeps the dominant patterns in its parameters and discards the rest, the same tradeoff a compressed memory makes.
Using SVD in a Titan-style memory module should, in principle:
- Compress information: SVD keeps the dominant singular directions of a weight update and drops the rest, similar to how a compressed memory keeps only the salient trace.
- Improve efficiency: low-rank updates are cheaper to store and apply than full-rank ones.
Review
Titan’s surprise-driven memory update
Titan (“Titans: Learning to Memorize at Test Time”, linked under Related Work below) updates a long-term memory module based on “surprise”: how unexpected new information is relative to what the memory already predicts. For an input token , Titan projects it into key-value pairs:
and defines an associative memory loss:
Using this loss, Titan updates a “surprise” momentum with forgetting:
and then updates the memory with a forgetting gate:
This lets Titan encode surprising information while a forgetting gate lets ordinary, expected inputs decay from memory.
Transformer² Singular Value Fine-tuning
Transformer² (“Transformer²: Self-Adaptive LLMs”, also linked below) fine-tunes a weight matrix by scaling its singular values rather than updating the full matrix. For a weight matrix , SVD gives:
where
- ,
- ,
- ,
with the rank of the decomposition.
This low-rank factorization cuts the number of trainable parameters while keeping the dominant structure of the weight matrix. Applied to a Titan-style memory module, the memory parameters become the set
which evolve over time as the model trains.
A design for integrating Titan and Transformer²
Combining Titan’s memory update with Transformer²’s low-rank parameterization gives a framework where an RL policy adjusts low-rank memory factors based on reward. Here is how that combination works out mathematically.
State, action, and transitions
At time , the state is:
where are the decomposed memory parameters.
The action adjusts these factors:
The transition updates the parameters:
which implies
The memory itself still updates via the surprise mechanism, applied with the next step’s decomposed parameters:
Policy and reward
A policy decides how to adjust the low-rank factors given the current state. The objective is to maximize expected cumulative reward:
where is the reward at time , reflecting how well the model’s memory helps it perform on the task at hand.
Policy gradient update
By the policy gradient theorem, the update for policy parameters is:
with
The policy update rule becomes:
Actions sampled from this policy update the SVD factors, which is the mechanism that ties memory adjustments to long-term reward instead of a fixed learning rule.
Conclusion
Combining Titan’s surprise-based memory update with Transformer²’s low-rank SVD parameterization gives a model where memory adjustments are learned via RL rather than fixed by a hand-tuned decay schedule. The math above is self-consistent, but it is a design, not a result: nothing here has been trained or benchmarked.
What would validate this
The next step is running it: fine-tune a small model with this policy, and check whether the learned updates actually improve task performance over the fixed forgetting-gate baseline. That experiment has not been run yet due to lack of stable GPU access; if it happens, the results belong in a follow-up post.
Related work
Paper:
Code: SakanaAI’s Transformer² GitHub Repository
Blog: SakanaAI’s Transformer² Blog Post
If you want to run an experiment like this yourself, the tooling below covers the data and observability layer for tracking training runs, and the newsletter is where the results of this one will ship if the GPU access comes through.
Next steps: scaling to production
If you take this into production, these are the pieces I would add first.
- Supabase Supabase is a hosted Postgres platform with authentication and storage built in. Postgres with pgvector for embeddings, so you do not run a separate vector store.
- Datadog Datadog aggregates metrics, logs, and traces for infrastructure monitoring. Traces and cost metrics across model calls, so latency and spend are visible per request.
- Vercel Vercel hosts frontend applications with a global edge network and CI/CD. Deploys the frontend and edge functions that sit in front of the model API.
Deploying generative AI models to production
Get the free playbook on shipping generative AI models to production.