Running local models is good now
writing custom tooling, such as a local HTTP [to] run the claude harness with a local model
OT, but to be clear, claude, as almost any other agent I've tried, can be configured with local models. Something like this will do:
There are, however, two issues with claude and most other harnesses I've tried:
1. telemetry or other outbound connections that cannot be disabled no matter how hard you try (which may be the reason you're running a proxy);
2. large system prompts that cannot be customized, which makes the tool impractical for local use unless you have a lot of resources. And even then, the prompt may not work well at all with the particular model you're using.
Regarding your point, it is not clear that the capabilities of large language models will continue to grow indefinitely, and not everything can be improved by throwing more hardware. The gap between what I can do locally and what I can do with a cloud system, although still present, might be significantly reduced.