AI-ANNE
GENERATIVE TRANSFORMER


(1) Initial Settings Architecture Hyperparameters


14 few (2) or many (24) dimensions

(2) Data Training Data


Vocabulary Size
Batches
Context Length
Attention Heads

(3) Training Training Hyperparameters


60 few (5) or many (80) words

10 few (3) or many (20) tokens

one (1) or four (4) attention heads

few (50) or many (1000) epochs

0.04 slow (0.01) or fast (1)
Multi-Head Attention (max. 2 visualized)
Loss Function (Cross-Entropy)

(4) Embeddings Mathematical Representations

TokenVector

(5) Sentence Generator Sampling Hyperparameters


Safety limit (if the end of the sentence is missing)

0.8 low (0.2) or high (2.0) creativity

5 few (1) or many (10) alternatives

0.90 low (0.1) or high (1.0) probability