Reading lists
- Code Is Not a File: Reading Graphs of SoftwareCode graphs model source code through syntax, control flow, data flow, dependencies, and semantics for analysis, navigation, security, and tooling5 sessionsLikes: 1Shares: 2Bookmarks: 1
- The Agent’s Second MindAI agent context management across conversations and workflows, including memory, retrieval, summarization, persistence, personalization, tool results, privacy, and reliability5 sessionsLikes: 0Shares: 0Bookmarks: 0
- From Next-Token Prediction to Trustworthy MachinesLarge language models: how they learn and generate text, why they seem intelligent, uses, limitations, reliability, and safety questions5 sessionsLikes: 0Shares: 0Bookmarks: 0
- When the Judge Becomes the BenchmarkFundamentals of using large language models as judges to evaluate AI-generated outputs, including rubrics, prompting, comparison, calibration, reliability, and bias5 sessionsLikes: 2Shares: 3Bookmarks: 1
- Beyond the Score: Evaluating AI Systems That ActMethods for evaluating AI systems and agents using benchmarks, tool-use tests, reliability and safety measures, human judgments, and real-world deployment5 sessionsLikes: 0Shares: 1Bookmarks: 1
- Beyond the Chatbot: Engineering Agentic SystemsAI agents and agentic architectures: components, planning, tool use, memory, coordination, execution, engineering, evaluation, deployment, limitations, and open research questions.5 sessionsLikes: 0Shares: 0Bookmarks: 0
- Memory Is a Context-Management ProblemAI agent memory and context management: architectures, retrieval, context windows, personalization, evaluation, implementation trade-offs, reliability, privacy, and open questions.5 sessionsLikes: 1Shares: 6Bookmarks: 1