Interview With Peter Bex

Peter Bex (website) is a CHICKEN scheme maintainer and professional Clojure developer. We got to know each other on IRC over a few months and discussed:

  • CHICKEN Scheme and its 6.0.0 release
  • Scheme’s community and standardization process
  • Postgres, the horrors of MySQL and (default) SQLite
  • Clojure
My ongoing core contributions mainly focus on the numerical tower code, keeping the in-core copy of the irregex library up-to-date with the upstream version (of which I’m a co-maintainer), squashing bugs and the odd security fix. I enjoy deep diving into odd corners of a code base to get a better understanding and then improve any weirdnesses I find. - His CHICKEN About

Beginnings

How’d you start computing?

I started computing when I got an old hand-me-down C64 which came with a “learn BASIC” book targeted at kids. Of course, I wanted to write my own computer games (which I never ended up doing). Later I got a real PC for my birthday and discovered my BASIC knowledge sort of transferred (with QBASIC in DOS).

Later, I learned about C, which I used for many years until at uni there was a course which taught Lisp (Scheme, really). We started with The Little Lisper/Schemer, and then SICP. I had a “functional programming course” before that which I almost flunked because it just didn’t make any sense (they used Concurrent Clean, because obviously teachers use their own language to teach regardless of its qualities - a common fault in academia). But the Little Schemer finally made it click. The teacher also showed how to implement objects using nothing but lambdas which I found awesome.

I actually studied AI before all the LLM bullshit. I much preferred the cleverness of the classical AI algorithms like A* search and genetic algorithms, but I haven’t really used them much in practice, only for my studies. I keep thinking I should use a GA for something, but no real use case so far. But even back then it was clear that neural networks were the future, though I found them boring because it’s just a bit of math (which I initially barely understood) and it’s basically a black box. I’m still grateful I did that because that’s the reason I came into contact with Lisp.

I might’ve looked into it myself at some point (having been curious about it as one of the “foundational languages”), but I might not have had sufficient gumption to really dig in. After uni I got a job where they were using Rails in those somewhat early days (2006), so I learned Ruby as well. I liked learning Ruby and thought it was cool to see the power of Rails, but later got very frustrated with it because Rails really has strong opinions and my programming style didn’t seem fully compatible with it.

CHICKEN & Scheme

Why CHICKEN?

After that course, I tried using Lisp for every personal project. I first started with Scheme48 but it wasn’t very practical (though very elegant). I remember running into problems with the image not being big enough, running out of memory. I never really liked PLT Scheme (now Racket), coming into first contact with it through DrRacket, which felt very sluggish to me, though I do think DrRacket’s a cool alternative view of what an IDE could do like when you hover over a variable and it shows you a line pointing to the origin of that variable.

CHICKEN was a practical and fast system, with a good community and acceptable license (I had quite a distaste for GPL at the time). CHICKEN is also one of two (as far as we know) implementations that use the Cheney on the MTA technique explained here. There were lots of rough edges, but I think wanting to address those is what enabled me to get so deep into the core. If everything is perfect, there’s not much to do, really!

So I started contributing to CHICKEN with some modest eggs (CHICKEN packages) at first, and eventually the core system. I mostly learned about Lisp internals by doing, mucking around with the CHICKEN core and trying things. I read Queinnec’s Lisp in Small Pieces, which deals with translation to C but leaves a lot undiscussed and distracts with OOP-heaviness. SICP has some good material. Then there’s Appel’s Compiling with Continuations, which is really short and to the point but still manages to be rather comprehensive; I love books like that.

In general, I just enjoy hacking on the core, even if I don’t have that much time to do so these days. I have to stress that I’m a slow learner. This understanding was a process of many many years, and there are still parts of the system I’m not that familiar with (although I know my way around enough to get up to speed if needed).

What do those Lisp books lack?

Compiling with Continuations has only a brief section on the runtime system, so it doesn’t go very deeply into e.g. garbage collection and data representation. And Lisp in Small Pieces doesn’t go into continuation passing style, IIRC. And its data representation isn’t very optimized. I like to blog about cool techniques that are undiscussed nowadays. For example implementing weak references and how to GC them efficiently. Other stuff I rarely see discussed is how to do FFI and cross-module optimizations, separate compilation and cross-compilation (which only a handful of Schemes even support) etc.

What’s Scheme to you?

In general, Scheme, to me, is a very clean language with a minimal core which facilitates experimentation. This is the fundamental tension of the standardization process - production-quality Schemes tend to grow in size, and there is value in standardizing that. But that also takes away the minimalism which makes experimental implementations possible. For example, Felix Winkelmann (CHICKEN’s original author, @Bunny351 talk about Scheme implementation) once started a Common Lisp (subset) implementation where he experimented a lot with types and flow analysis, resulting in CHICKEN’s type stuff.

What are your thoughts on the overall Scheme ecosystem, R7RS, SRFIs vs. implementation-specific libraries etc.? A few implementations don’t seem to care anymore.

I think the split of R7RS into a small and large language was the right thing to do, as R6RS was reviled by minimalists and found lacking by maximalists. In general, I’m a bit sad for R7RS - if some of the bigger community Schemes are essentially completely ignoring it, they’re doing something wrong IMO. Many, maybe even most Scheme implementations are essentially one-man shows. I suppose CHICKEN is also turning that way again, since we have lost quite a few contributors (mostly due to changing life situations) and it’s hard to attract new ones.

At the same time, the R7RS “large” project sort of went off the deep end, doing its thing without really caring about community buy-in. The churn of the R7RS large is also a bit too fast to keep up with. They’ve pumped out tons of SRFIs in a few years (actually, looking back at it right now it doesn’t appear like it’s that much, but it is certainly a lot faster than SRFIs used to be.) The SRFI process is open to submissions from literally anyone, for better or worse. I’ve noticed there have been a few new contributors to CHICKEN who submitted implementations for some of the newer SRFIs, so that’s good and I’m happy at least some people are bothering to do this.

CHICKEN 6.0 is base R7RS, but older CHICKEN code will keep working. We still support the old module syntax (which is the “native” one). The R7RS library declaration is essentially syntactic sugar for the core module syntax. R7RS-small is almost fully backwards compatible with R5RS, so there is no conflict there. Porting an egg to CHICKEN 6 usually requires a few small adjustments because some (non R5RS) identifiers moved around between modules to better fit the R7RS style.

What’s CHICKEN’s development process like?

CHICKEN 4 was “hygienic CHICKEN”, which introduced the module system (and required overhauling the expander). This was all Felix, requiring a lot of deep internal knowledge about how macros interact with modules etc.

CHICKEN 5 was a community effort through and through. It was mostly a sanity and cleanup release where we did a massive reorganization of the modules (what lives where) to make it logical and matching R7RS a bit better. We discussed this during an IRL meetup and continued for the months after. (Community is a strong advantage of CHICKEN!) We also added a numeric tower.

CHICKEN 6 made it backwards incompatible, resulting in a new major release. It’s basically cut off from Felix’s branch to make UTF-8 handling sane and consistent (like Python 2 -> 3, but way less disruptive.) There’s a strict separation between strings and bytevectors, with changes to ports and other I/O as well. We took the opportunity to integrate the R7RS egg into core so it’s more “native”. Strings in the FFI should be more efficient because there’s no needless copying anymore.

CHICKEN 6.0.0 was held back by a bug causing heap corruptions (do view the patch’s description!). We had been looking in the completely wrong spot. These heap corruptions appeared somewhat randomly, but only with the CHICKEN wiki server, not with the plain web server serving simple files or even the entire Awful framework. We strongly suspected the Subversion client library (which the wiki uses as a backing store for the content), and we’d found other issues in there as well (it’s kind of hairy callback-heavy code due to the design of libsvn).

The wiki is a rather small program and the rest of the web stack seemed to be fine, so we suspected the svn client lib. But I whittled down the code of the wiki to almost nothing and it was still failing. When I commented out the URI normalization code (you get redirected when opening a page that’s behind a symlink, so as to get a canonical URL that points to the original file) it suddenly stopped crashing!

That normalization code didn’t appear to do all that much, so we quickly pinpointed it to be read-symbolic-link. A quick glance at the code in the core system confirmed it was totally borked because of a change made for CHICKEN 6’s UTF-8 transition.

Now that 6.0.0 has come out, what’s next?

Regarding the goals for CHICKEN, there are several things I’d like to work on. One idea I had is to teach the compiler about unsafe intrinsics using a “prelude”. Because Scheme is a safe but dynamically typed language, there’s some overhead in the intrinsics, say if you call “car” on a non-pair, it throws an exception. If the compiler can deduce that the object you pass in must be a pair (maybe because you checked it before with pair?, or called car or cdr on it before), it replaces the call to an unsafe, unchecked version. But this is all very ad-hoc, and not extensible by the user. My idea was to have a separate definition which splits the unsafe operation from the “typechecking prelude”, which can be inlined at the call site. This way, if multiple checks need to be done, it’s not all or nothing. We can elide the unnecessary checks and only do the necessary ones which might extend to user code, too.

Another idea relates to the way we handle dates and times - we have some stuff in core to access the POSIX functions but it’s messy and (IMO) mostly unusable. The alternative is SRFI-19, which is a beast because it has support for multiple calendar systems, localization etc. Might be nice to have something minimal (maybe English only) in core, so you have a common type that gets used everywhere (handy when sharing objects between libraries without building in a big dependency on SRFI-19). You can then use it for parsing timestamps in common protocols, say.

Postgres & SQLite

What domains do you like or know the most about?

  • Web stuff: I maintain the HTTP and URI implementations for CHICKEN
  • Some CHICKEN internals: GC, macro expander and Irregex implementation
  • Performance optimizations: though not an expert, I’ve done quite a bit and always thoroughly enjoy it
  • Postgres: Although I haven’t gone deep into the internals, I’m typically the go-to guy for (Postgre)SQL questions in companies I’ve worked at. Funny, because I initially flunked the DB/SQL course at uni and didn’t grok SQL at all
  • Distributed systems: though I’ve worked on them for 6 years at work, you’ll want to avoid them like the plague if at all possible. It can be hard to reason about the behaviour of the system at large, and you can’t really abstract it away

I’m trying really hard to think of something I’m truly excited about. The biggest positive I see right now is the push for digital sovereignty. I sincerely hope this will change how people deploy tech, maybe in a more mindful manner. More open source, less dependence on foreign (and hopefully big tech in general) products. But vested interests and inertia will be hard to overcome and really bum me out.

Why Postgres?

I properly learned about DBs at a calendar startup using Rails. We had instantiated repeated events in the DB and the event would sometimes need to be updated. At first, we were fetching models in a loop and updating them one by one, excruciatingly slow. We eventually discovered the bulk update (I think you even had to call into the DB directly because Rails didn’t offer that at the time). That made everything click for me - the importance of performance and the usefulness of SQL. At the time, MySQL was still Rails’ default and I got into MySQL character set hell a few times for a CMS we used. Later, another Rails project required such massive amounts of data to be stored (computational fluid dynamics simulation) that MySQL simply crashed every time I tried a bulk import.

Looking into alternatives, I found Postgres handled it without any problems. When I learned that Postgres DOESN’T have any of those braindead misfeatures MySQL has. For instance, UTF-8 characters get verified on storage so you can’t get into character set hell like in MySQL so easily, and it actually allows DDL statements in a transaction, so you get transactional migrations that apply atomically. That was a real eye opener at the time. I was sold!

Postgres is a lot more regular and well-behaved on basically anything, and it has no strange limits (e.g. in MySQL, you can’t even put an index on text columns with indeterminate lengths). MySQL allows you to store an empty string in any non-nullable enum column. Makes no goddamn sense to me! The DB is full of footguns like that. I should stop ranting - talking about MySQL really makes my blood boil. I like to rely on more “advanced” features like LISTEN/NOTIFY, array storage, window functions, CTEs etc. I’ve found that stored procedures don’t really work that well - “real code” is more flexible as it doesn’t require finicky migrations to keep in sync.

I’m not a big fan of JSON in my relational DBs, but I have been known to use it when storing arbitrary data or actually putting JSON results (from APIs or other stuff) in the DB.

We use DataScript on the client via ClojureScript, but I don’t really grok it and don’t touch that part of the code often enough that it really sticks, so every time I have to deal with it it’s an exercise in frustration.

Even though I grok it nowadays, SQL is a really badly designed language. I’ve seen several projects that try to come up with a better query language, but I’ve given up hope that they will succeed as SQL is too entrenched to get rid of.

Why not SQLite?

I’ve used it a handful of times. The experience was always mostly one of frustration. It feels a lot like MySQL, with unsafe and stupid defaults.

IIRC it’s value-typed and (by default) doesn’t check types, so the type of a column is basically completely ignored. And you can’t alter a column, IIRC (or maybe only a few changes). I even remember a version (maybe WebSQL?) where you can’t even DROP a column. Anyway, it’s not worth any brain cycles for me to deal with that shit.

On REPLs

What do you think of Clojure?

Clojure’s influence is strong on things like Carp and Janet. I wrote about my impressions about Clojure on my blog, but in a nutshell I don’t like the “everything is a map” approach and nil punning really turns me off as it makes bugs harder to find. I do like the fact that it has mostly purely functional data structures. I’m not sure I like the syntactic “heaviness” of Clojure - things like [] for vectors and {} for maps. The lack of cons cells is also a bit weird but I see how it simplifies list handling code (even though Clojure does not typically deal with lists much!) Most importantly, it revitalised the interest in Lisps!

In your article on Clojure:

never fully bought into the REPL style of developing. Sure, I experiment all the time in the REPL to try out a new API design or to quickly iterate on some function I’m writing, but my general development style tends more towards the “save and then run the test suite from an xterm”.

Normally, we hear such things from people who think using the REPL means typing into the little terminal box instead of sending code from files into the REPL with a hotkey, so it surprised me to read it from a veteran.

I do consider the REPL an essential tool for experimentation and debugging. But I struggle to keep track of what’s running in the system versus what I see in my editor. With Clojure, you can’t do without the REPL because it’s so doggone slow to start up that it would be impossible to just run something on the CLI over and over. So at work, I spend 100% of my time with the CIDER REPL. I do find myself closing and reconnecting several times a day though, because I can no longer trust the REPL state matches my editor buffers. One thing that gets me every time is if I delete a test from my buffer and then re-run the entire suite, it’s still there in the REPL (obviously), same thing with multimethod implementations. I really couldn’t live without a REPL.

What do you think of snapshot testing?

Snapshot testing’s an interesting approach. I think we have a few testcases in CHICKEN where we do something like that - we run the compiler and capture the output of the compiler (which is mostly type warnings) and check that it hasn’t changed with a simple diff on the output and expected output. I’ve also used something like this for great effect while refactoring and optimizing code - simply keep reference output in a file. For instance, in one case when working on a project which had no test suite, I used pg_dump to dump an “output table” and then went to town on the codebase, knowing I would immediately see if my optimized algorithm differed from the original. I also used this in another inherited project without test cases to refactor.

Real Life

What makes you happy?

I’ve mentioned before that I really enjoy performance optimizing code, but I also really enjoy refactoring and investigating vulnerabilities. Systems that are understandable and hackable make me happy.

Outside of programming (which more often frustrates me than makes me happy TBH), my family makes me happy. It’s a great source of joy to just relax and be with my wife and children. I’ve been trying to get back some balance in life, spending more time AFK, and I pay attention to my health a bit more as well.

I’m not super young anymore (43) and reading about age-related issues like sarcopenia made me realize we really tend to neglect our bodies with our sedentary lifestyle, especially us programmers. My mother has osteoporosis and I see how she struggles just doing basic things. I don’t want that for myself, so I’ve picked up weight lifting as a way to combat those age-related issues, so I can become old in a healthy way. The prognosis is for most of us to live up to 90 or so by now, so I’m not even at the halfway point. But these issues start cropping up at around 50-60.

How do you approach raising kids?

I’m still getting my bearings TBH. My kids are only 3 (going on 4) and 15 months. Raising kids is probably the hardest thing I’ve ever done.

I don’t really know yet if I want to teach them programming. I definitely want to raise them tech-sceptical, when they’re big enough, to teach them the dangers of social media and the importance of privacy. If they show an interest, obviously I’d teach them Scheme. It’s the perfect language for teaching! But maybe something like Logo first.


Appendix from #scheme

I spent some time in #scheme on IRC, even attending an R7RS meeting.

People argue Common Lisp’s many implementations are a strength, but Go basically went to only 1. Besides portability to other platforms, why have so many Scheme implementations?

retropikzel: I just think they’re neat. And I want more of them. Some day maybe I’ll make one myself Hopefully also Scheme can offer an ecosystem to tap into. If someone is implementing a lisp, we could get more schemers by saying “Hey add RnRS support and here is libraries you can use”

Wolfgang Corcoran-Mathe: Personally, I do not like the situation the Haskell language is in. (One de-facto-standard implementation that’s much too big to fail.) I appreciate having a lot of implementations to choose from. The implementations have different goals. e.g. chibi is aimed at low memory usage, CHICKEN at integration with C, Guile at being a huge ball of mud, I mean, flexible …

phm: In my opinion, a single implementation can cover many use cases but the implementation itself can become very complex in doing so. Having multiple implementations cover a similar set of use-cases is easier because most people only need to cover a few use-cases. There are also use-cases that are very difficult to reconcile in one thing, like embeddable (Chibi Scheme) vs featureful (Racket). I can go into the code of Chibi, Gambit, and CHICKEN and have an idea what’s going on. It would take me a lot longer to do so with LLVM, even though I have written C(++) for much, much longer than I have Scheme. Even Chez and Racket’s source code isn’t that bad, and they’re also very fancy.

Which implementations do you normally use?

Wolfgang Corcoran-Mathe: Chez, mostly, these days. But I’ve been using Gauche a lot for R7RS stuff.

retropikzel: I test my r7rs only code usually on chibi, chicken, gauche, kawa, mosh, racket, sagittarius, skint, stklos, tr7 and ypsilon. Also using akku-r7rs, I often test on chezscheme, ikarus, ironscheme, racket(r6rs), sagittarius(r6rs), ypsilon(r6rs). Akku mirrors snow-fort packages so I want to confirm it works. Then there are implementations I test less often, cyclone, gambit, guile, larceny, loko, meevax and mit-scheme. Mostly because they dont have good enough r7rs support to not be annoying to work with sometimes. A combination of those will give you enough of combined error messages to make debugging easier and if something works on those it will probably work on most.

What do you actually want scheme to be? I know a few e.g. write Clojure at work but prefer Scheme. I know traditionally there’s the “elegant minimalism” vs. “we need to be able to write actual software” dichotomy…

aeth: There’s always a few directions pulling at Scheme designs. Let’s copy what CL is doing (but make it fit Scheme’s more minimal sensibilities), let’s reject what CL is doing because we’re not CL, let’s do what’s trendy in FP. There’s kind of three tensions particularly Scheme-ish because Scheme is academic. Beginners/students, CS research, the software industry. Separately, scripting vs. application is a big tension. CL is all about being an all CL application, most other Lisps are limited by their major implementation (e.g. elisp is obviously a scripting language). Scheme seems to split along the lines of embedding a little scripting language vs being its own standalone thing.

Wolfgang Corcoran-Mathe: I think we’d be fooling ourselves if we adopted the goal of making Scheme trendy. The language has a tradition of The Right Thing, which I would like to continue.

phm: I want Standard Scheme to be something expressive enough where you can write portable code for many implementations, such that you can choose one to write your practical software.

phm: I like thinking of Scheme as a low-level functional language. Which sounds weird to say, but in a way that a low level language is “close to the hardware,” Scheme is close to its theoretical underpinnings, with call/cc and all that. If you want an actual low-level Scheme, CHICKEN and Gambit are both Scheme -> C compilers that allow you to splice in C code anywhere. So you can do some fun stuff with it.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论