ICA Communities in Kimi K3 Token Embeddings

This is a small application of our ICA Lens pipeline. Kimi K3 has almost 100 layers, and I do not have enough memory to load all of its weights, but I can analyze its token-embedding matrix by itself.

Token embeddings are often projected into a two-dimensional cloud with PCA or t-SNE. Here I instead fit FastICA, selected the 1,000 most non-Gaussian components from 7,168, and grouped them into nine communities. For each component I retained its ten highest-scoring tokens on the dominant sign. Together, these 10,000 token entries give a compact view of relationships among the embedding directions.

The FastICA fit took 5 minutes 30 seconds using FastICA_torch. The heatmap below shows absolute cosine similarity, ordered by community. Click a column to inspect its component, community, objective, and top tokens.

The community labels are provisional summaries of the top tokens, not ground-truth feature names. Tokens can also appear in more than one component. Even with those limitations, ICA exposes structure that is difficult to see in a single two-dimensional projection.