Q&A #88 (2026-10-02)

In each Q&A video, I answer questions from the comments on the previous Q&A video, which can be from any part of the course.

As an additional note, for those who are interested in the Loop Stream Detector from question #6 and don’t mind doing some additional reading, there is a security exploit called GADGETSPINNER you might want to look at. The writeup is helpful for learning about the particulars if, like me, you don’t have an Intel CPU with an LSD to play around with yourself. Of potential interest for people focused on optimization are the LSD’s dependence on the branches staying the same each iteration, as well as the total uop size of loop that it can handle.

The questions addressed in this video are:

  • [00:02] “Some time ago you were investigating some hardware for low level programming (Raspberry Pi, Beagle Board). Do you have any suggestions for someone who wants to try some old school game development with this kind of limited hardware? (Emulating old systems doesn’t really do it for me)”
  • [02:42] “For some time now I am banging my head against the wall regarding timesteps, framepacing and vsync. I think I understand the gist of how it works in general (frame queues, presentation modes, composition, etc.), but when I try concrete implementations there’s always some kind of catch to work around, which makes it wonky. Like GetFrameStatistics works pretty well, but breaks on multi monitor setups. Manual timing and waiting seems to be too inprecise over time and gives weird results with vsync. What would be the (imaginary, magic) api call needed, to get a metric which solves this once and for all?”
  • [19:38] “Has the situation regarding the 30 million line problem improved by now? More specifically, I wanted to know. If it’s feasible for someone to make their own mini os for their own mini game that can run on multiple pcs (only targetting modern hardware but also supporting gpu rendering for newer amd cards).”
  • [23:59] “Hi, more on memory mapped files: the impression I got from the course is that they’re quite bad at nearly everything - you don’t have a lot of control over page faults and when / how io happens. But many fast databases, say lmdb, uses memory mapped files and have excellent performance. Could you maybe explain a little bit more nuance on that side, such as where they have a big advantage and should be used? Or maybe you believe that lmdb authors could have achieved similar performance with less effort?”
  • [30:02] “I hear people usually say ‘nah, everything been written before us anyway’ or something like that, meaning that you can’t really write anything new, just rebuild/repackage/reinvent the wheel. and you clearly have a different view here. so could you please describe in more details why you think differently (or thought, if something changed)”
  • [35:52] “I just tried running SpacingLoop with Rocket Lake on uops.info, and I noticed that it shows LSD instead of DSB. I looked it up and found that LSD means Loop Stream Detector. Could you explain how it works?”
  • [40:25] “A question about those masked instructions: we can save some instructions and work by instead of separately computing two values and then recombining them we just do different computations on the same register with different masks: first computation touches a subset, second computation touches a subset. But doesn’t it create a serial chain? Now hardware can’t run both versions in parallel, which may be more efficient for longer chains, but has to do each operation after another? Or is it that smart that it tracks serial dependencies based on masks instead of registers?”
  • [48:43] “Tim Davis (sparse matrix algorithms) pointed out that modern machines that rely heavily on caches are actually very similar to theoretical turing machines that have to pay the cost of a navigating head.I thought that was another instance of John Backus delaying the compiler technology besides him ignoring the SSA construction from Hoar. as he famously wrote ‘Can programming be liberated from Von Neumann style’. Which turns out not to be a style but a much more grounded method. and funny enough that SSA is just the recovery of functional benefits within this overarching method. I would be very happy if I can hear your thoughts around that, even if just to point out that I am hallucinating.”
  • [51:35] “In light of recent videos on removing predication, I’d like to ask about the conditional move instruction. Is this something that’s considered a good thing, if you can’t remove predication otherwise?”
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论