Projects
4 projects
local session · saved
01 Gamma Recurrence human + tutor
Gamma Function · Gamma Recurrence3 objects

Don’t finish it. Mark the exact place my reasoning breaks.

Γ(9/2) = ∫₀∞ x⁷ᐟ²e⁻ˣ dx = [−x⁷ᐟ²e⁻ˣ]₀∞ − (7/2)Γ(7/2)

Gamma Function · Gamma Density3 objects
ga(x)=xa1exΓ(a),0ga(x)dx=1g_a(x)=\frac{x^{a-1}e^{-x}}{\Gamma(a)},\qquad\int_0^\infty g_a(x)\,dx=1
Tutor
normalised gamma density · total area 1

Gamma densitya = 4.5 · b = 6

ga(x)=xa1exΓ(a)g_a(x)=\dfrac{x^{a-1}e^{-x}}{\Gamma(a)}
w₁ = 0.290w₂ = 0.505w₃ = 0.205P(X ≤ 6) = 0.787mode a − 1 = 3.50481216
Γ(4.5) = 11.6317 · edges 3.15, 6.08
w₁[0, 3.15)w₂[3.15, 6.08)w₃[6.08, )Σ
probability mass0.2900.5050.2051.000
log mass ℓⱼ-1.236-0.684-1.585
softmax(ℓ)ⱼ0.2900.5050.2051.000
the final bin owns the tailw₃ = 1 − w₁ − w₂, so the masses sum to exactly one.log-masses are the logits the attention head starts from.
wj=bj1bjga(x)dx,j=logwj,softmax()j=wjw_j=\int_{b_{j-1}}^{b_j}g_a(x)\,dx,\qquad \ell_j=\log w_j,\qquad \operatorname{softmax}(\ell)_j=w_j
Tiny Transformer · Attention2 objects

One head, two-dimensional embeddings. Every value on this card is recomputed from the matrices you edit.

TINY TRANSFORMER · ONE HEAD · 2-D EMBEDDINGS

Attention is a weighted sum

softmax over (q·kⱼ/√2 + log wⱼ) / TT = 1 · query target · target x₁
KEYS k = W_K eQUERY q = W_Q e53°14°24°kkkc = Σ αⱼ vⱼq = W_Q e₂e · x₀e · x₁e · target
SOFTMAX WEIGHTS αderivedΣ α = 1.000
α · x₀0.267
α · x₁0.534
α · target0.199
EMBEDDINGS e · CLICK A TOKEN TO QUERYeditable
target
query
WQ
WK
WV
TEMPERATURE T
TARGET
SCORES → SOFTMAXderived
tokenq·kⱼ/√2+ log wⱼ÷ Tαⱼ
x₀0.155-1.236-1.0810.267
x₁0.298-0.684-0.3860.534
target0.212-1.585-1.3720.199
CONTEXT c derived[0.103, 0.497]
LOSS derived1.006−log p(target), p = 0.366
NEXT TOKEN pderived
x₀0.307
x₁0.366
target0.327
weights, embeddings, temperature and the query/target choice are yours · everything tagged derived is recomputed on every edit
Tutor
Tiny Transformer · Gradient Step2 objects

A tiny transformer training step: one numerical gradient on the visible parameters, never frontier-model training.

Tutor
TINY TRANSFORMER · GRADIENT STEP

One honest training step

STEP0η
RATE η editableline search starts here, halves until it improves
TARGET · QUERYx₁query target · shared with the attention card
LOSS derived1.006awaiting a step
P(TARGET) derived0.366awaiting a step
OUTPUT p derivedΣ p = 1.000
x₀0.307
x₁0.366
target0.327
LOSS · 1 point1.006
1.006
P(TARGET) · 1 point0.366
0.366
WEIGHTS BEFORE → AFTER · SINCE STEP 0at step zero
W_Q
0.820.82-0.18-0.180.220.220.740.74
W_K
0.680.680.280.28-0.16-0.160.760.76
W_V
0.640.64-0.24-0.240.180.180.70.7
central finite differences on every visible parameter · deterministic backtracking from η · a step is committed only when the loss fell and p(target) rose
Olympiad Geometry · Barycentric Coordinates2 objects
P=αA+βB+γC,α+β+γ=1P=\alpha A+\beta B+\gamma C,\qquad \alpha+\beta+\gamma=1
Tutor
Barycentric coordinates · Ceva live

P = αA + βB + γC

Σ = 1.000weights = attention α
DEFα = 0.27β = 0.53γ = 0.20ADouble-click to renameBDouble-click to renameCDouble-click to renameP
signed subareas[-13775, -27610, -10294]÷ -51680 =[0.267, 0.534, 0.199]

Dragging a vertex moves P by the same weights: the affine combination is the invariant.

Ceva · BD/DC × CE/EA × AF/FB = 0.37 × 1.34 × 2.00 = 1.000

Tutor
Olympiad Geometry · Spiral Similarity4 objects

Two circles tangent at O. Drag A.

OAhOA=OBhOB=OChOC=0.58\frac{OA_h}{OA}=\frac{OB_h}{OB}=\frac{OC_h}{OC}=0.58
Tutor

Drag A, B, C or O. Mapped points, tangent circles, ratios and equal angles recompute.

Dynamic geometry

Homothety · Spiral similarity

10 points · 14 marksmove · click to select
ODouble-click to renameADouble-click to renameBDouble-click to renameCDouble-click to renameAₕDouble-click to renameBₕDouble-click to renameCₕDouble-click to renameA′Double-click to renameB′Double-click to renameC′Double-click to rename
OAₕ / OA0.580OBₕ / OB0.580OCₕ / OC0.580drag A, B, C or O · the two circles stay tangent at O
Tutor
Simplex and Partitions · Simplex2 objects
P=αA+βB+γC+δD,α+β+γ+δ=1P=\alpha A+\beta B+\gamma C+\delta D,\qquad \alpha+\beta+\gamma+\delta=1
Tutor
4-weight probability simplex · perspective projection

Tetrahedral probability

P [0.24 : 0.41 : 0.17 : 0.18]·orbit -0.32 / 0.49
ABCP [0.24 : 0.41 : 0.17 : 0.18]BCAD
56lattice points · L₃(5) = C(8, 3)Pascal 35 + 21 = 56
N
Simplex and Partitions · Integer Partitions2 objects
P(q)=m1(1qm)1=n0p(n)qnP(q)=\prod_{m\ge1}(1-q^m)^{-1}=\sum_{n\ge0}p(n)\,q^n
Tutor
Integer partitions · finite Euler product

Lattice points become partition coefficients

lattice point · N = 5(1, 2, 1, 1)one of L₃(5) = 56 tuples
partition of 5 · ≤ 4 parts2 + 1 + 1 + 156 tuples ≠ p(5) partitions
finite Euler product, m ≤ 14m=11411qm=n14p(n)qn+\prod_{m=1}^{14}\frac{1}{1-q^{m}}=\sum_{n\le 14} p(n)\,q^{n}+\cdotscoefficients below are computed here
Ferrers · 5 cells
p(9) = 30n ≡ 4 (mod 5) · theorem lane
n ≡ 0 (mod 5)
p(0)11p(5)72p(10)422
n ≡ 1 (mod 5)
p(1)11p(6)111p(11)561
n ≡ 2 (mod 5)
p(2)22p(7)150p(12)772
n ≡ 3 (mod 5)
p(3)33p(8)222p(13)1011
n ≡ 4 (mod 5)
p(4)50p(9)300p(14)1350
Tutor
selectright-drag to pan
100%