Sources & Further Reading
iSources checked: September 2026
Sources reviewed in September 2026: links checked for availability and extended with references for the current state. As the AI landscape evolves quickly, some details may have changed since then.
Sources
-
Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762. https://arxiv.org/abs/1706.03762
-
Mikolov, T. et al. (2013). Efficient Estimation of Word Representations in Vector Space. arXiv:1301.3781. https://arxiv.org/abs/1301.3781
-
Muennighoff, N. et al. (2023). MTEB: Massive Text Embedding Benchmark. arXiv:2210.07316. https://arxiv.org/abs/2210.07316
-
OpenAI (2024). Text Embedding 3 – Documentation. https://developers.openai.com/api/docs/guides/embeddings
-
Hendrycks, D. et al. (2021). Measuring Massive Multitask Language Understanding (MMLU). arXiv:2009.03300. https://arxiv.org/abs/2009.03300
-
Chen, M. et al. (2021). Evaluating Large Language Models Trained on Code (HumanEval). arXiv:2107.03374. https://arxiv.org/abs/2107.03374
-
Jimenez, C. E. et al. (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv:2310.06770. https://arxiv.org/abs/2310.06770
-
Artificial Analysis – LLM Performance Comparisons. https://artificialanalysis.ai/
-
SWE-bench (2026). SWE-bench Verified leaderboard. Reference for the coding scores cited in this module. https://www.swebench.com/ · https://llm-stats.com/benchmarks/swe-bench-verified
-
Anthropic (2026). Claude – model overview and pricing. Reference for model names, context windows, and token prices. https://platform.claude.com/docs/en/about-claude/pricing
-
OpenAI (2026). Models – API documentation. Reference for the GPT-5.6 variants Sol, Terra, and Luna. https://developers.openai.com/api/docs/models/
-
Cho, A. et al. (2024). Transformer Explainer: Learning LLM Transformers with Interactive Visual Explanation and Experimentation. arXiv:2408.04619. The basis for the pipeline and attention visualisations in the "From Prompt to Answer" section. https://arxiv.org/abs/2408.04619
Further Reading
Transformer Architecture
- Jay Alammar – The Illustrated Transformer: The best visual explanation of the Transformer architecture. https://jalammar.github.io/illustrated-transformer/
- Transformer Explainer (Polo Club, Georgia Tech): A real GPT-2 small running in the browser -- with all intermediate results including the actual Q/K/V matrices and adjustable temperature, top-k and top-p sampling. https://poloclub.github.io/transformer-explainer/
Embeddings & Benchmarks
- Hugging Face – MTEB Leaderboard: Up-to-date ranking of embedding models across various tasks. https://huggingface.co/spaces/mteb/leaderboard
- SWE-bench Leaderboard: Ranking of code generation models on real GitHub issues. https://www.swebench.com/
Deep Dives
- Lilian Weng – Attention? Attention!: Detailed blog post on attention mechanisms. https://lilianweng.github.io/posts/2018-06-24-attention/
- 3Blue1Brown – Attention in Transformers: Visual explanation of the attention mechanism. https://www.youtube.com/watch?v=eMlx5fFNoYc