Guys... OpenAI API on VLLM and Llamacpp already supported grammar enforcer... (AKA JEV)

preview.redd.it/tcdyumkf3zsh1.png If you want to try JEV like generation, or what we could call an already fucking exist, zero-shot, training-free classifier , your LLM already supports it. The model running on the very computer you host does not need any server-side modification. The feature is called structured output, and the underlying idea is grammar-constrained generation or a grammar enforcer. Under the hood, vLLM supports multiple structured output backends, such as XGrammar and lm-format-enforcer, while llama.cpp uses GBNF. Basically, the decoder constrains the LLM so it can only generate tokens that are valid under the specified grammar or schema. For vLLM: docs.vllm.ai/en/v0.8.2/features/structured_outputs.html For llama.cpp, structured output is wrapped in a different JSON request-body format, or you can use GBNF directly. vLLM: structured_outputs └── json └── {schema} llama.cpp: json_schema └── {schema} This is example of that schema in vllm of product sentiment analysis, which roughly mapped to most of jev use cases: SCHEMA = { "type": "object", "properties": { "sentiment": { "type": "string", "enum": ["negative", "neutral", "positive"], }, "score": { "type": "number", "minimum": -1.0, "maximum": 1.0, }, "value": { "type": "string", }, }, "required": ["sentiment", "score", "value"], "additionalProperties": False, } def classify(news: str) -> dict: payload = { "model": MODEL, "messages": [ { "role": "system", "content": SYSTEM_PROMPT, }, { "role": "user", "content": news, }, ], "temperature": 0, "max_tokens": 512, "chat_template_kwargs": { "enable_thinking": False, }, "response_format": { "type": "json_schema", "json_schema": { "name": "news_sentiment", "strict": True, "schema": SCHEMA, }, }, } resp = requests.post( ENDPOINT, json=payload, timeout=30, ) Look, I think JEV and what it brings to the community as a refresher on already great classifier-style workflows is a plus for me. I just want to ground the discussion in the fact that this already exists, and you do not need a custom model just to study or experiment with training-free classification. I am very familiar with this because it is part of my profession in lakehouse platforms. Basically, we use 1B to 4B models to ingest unstructured data such as images or documents, then extract structured information such as place, time, sentiment, entities, and so on. Why? Because working with well-formatted SQL data is much less of a pain in the ass than repeatedly querying raw unstructured content through an Elastic/OpenSearch index.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论