The scaling laws hold that a language model grows more capable with more parameters and more training data, and Mixture-of-Experts (MoE) architectures have ridden these laws to remarkable results, activating only a fraction of an enormous stored parameter…
1. What should be our Bayesian priors on von Neumann probes? 2. AI book mirrors. 3. Why power in Spain is so cheap. 4. The comments on Mick West here are pretty tough.…
A Conversation With Murray Cantor (Part 2) In Part 1, Murray Cantor reframed technical debt as economic liability and argued that uncertainty is not an annoyance to be minimized, but the core feature of IT investment.…