<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1080/00207543.2014.942012</article-id>
      <title-group>
        <article-title>Copilot: Towards Integrating Large Language Models and Constraints</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Philipp Kogler</string-name>
          <email>philipp.kogler@siemens.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wei Chen</string-name>
          <email>chen.wei@siemens.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andreas Falkner</string-name>
          <email>andreas.a.falkner@siemens.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alois Haselböck</string-name>
          <email>alois.haselboeck@siemens.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefan Wallner</string-name>
          <email>stefan.wallner@siemens.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Product Configuration</institution>
          ,
          <addr-line>Constraints, Feature Models, Large Language Models, Copilot</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Siemens AG Österreich</institution>
          ,
          <addr-line>Siemensstraße 90, 1210 Wien</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>We utilize a pre-trained Large Language Model</institution>
          ,
          <addr-line>LLM</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <abstract>
        <p>A product configurator enables the configuration of a customizable product while constraining possible variations. Users typically interact with a product configurator via a graphical user interface. A complex product can be composed of components and parameters that are not easily understandable for non-experts which can prevent them from efectively configuring the product. In this paper, we propose a configuration copilot, an interactive chat-based interface that allows users to iteratively configure a product by describing their requirements in natural language. Our framework leverages the Natural Language Processing (NLP) capabilities of advanced pre-trained Large Language Models (LLMs) alongside the robustness of constraint-based product configurators. We introduce a technical architecture that accurately formalizes constraints from natural language inputs, identifies valid product configurations based on a defined product line and specified constraints using a constraint solver, and communicates the resulting product configurations back to the end user in natural language. We demonstrate and evaluate the configuration copilot on two use-cases: The configuration of the GoPhone feature model (Boolean feature assignments), and the configuration of a metro wagon (more general configuration parameters).</p>
      </abstract>
      <kwd-group>
        <kwd>of reliability</kwd>
        <kwd>guaranteed correctness</kwd>
        <kwd>domain-specific</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Product configuration involves creating customized
products from predefined components while satisfying
constraints that limit configurable parameters and possible
combinations [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. A product configurator is a software
tool that allows users to configure a product, commonly
through a graphical user interface and often in a
websign plays a major role in the development of a product
configurator but is often overlooked [
tion is especially relevant when complex products are
configured by non-expert users. The meaning of
configurable components and parameters may not be obvious
which prompts a need for explanation and introduces a
learning curve.
uct configurators, we propose a configuration copilot
that ofers a text-based chat interface. Uninformed users
shall be able to describe their requirements in natural
language without knowledge of the concrete parameters to
set and components to select. The copilot shall then
conifgure the product and respond with a valid configuration
      </p>
      <sec id="sec-2-1">
        <title>In this paper, we first describe LLMs and constraint</title>
        <p>in Section 3. We detail the technical architecture of the
configuration copilot in Section</p>
      </sec>
      <sec id="sec-2-2">
        <title>4, and present an eval</title>
        <p>uation based on the two use-cases of configuring the</p>
      </sec>
      <sec id="sec-2-3">
        <title>GoPhone feature model and a metro wagon in Section 5.</title>
      </sec>
      <sec id="sec-2-4">
        <title>We conclude the paper with a summary, a limitation statement, and future work in Section 6.</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Background</title>
      <sec id="sec-3-1">
        <title>2.1. Large Language Models</title>
        <sec id="sec-3-1-1">
          <title>Pre-training task-agnostic aspects of natural language</title>
          <p>processing (NLP) tasks is a central concept of LLMs.</p>
          <p>The Transformer architecture enables this approach
models are able to capture complex patterns and long- to the constraint solver. The results of the solver are
range-dependencies in texts through the multi-head presented on the GUI, and the user can vary or refine
self-attention mechanism. Compared to previous state- her/his input specification and the solver is called again.
of-the-art models such as recurrent neural networks To design and implement a configurator GUI can be
(RNNs) or long short-term memory networks (LSTMs) a challenging task, because the possible interactions are
a performance improvement in various NLP tasks is ob- diverse, like collecting the requirements, reporting
inserved [4, 5]. valid constellations, representing a solution, showing a</p>
          <p>Decoder-only models are a subclass of Transformer- performance value of a solution, etc. In addition, every
based architectures and are primarily used for sequence- modification of the product (line) requires a review and
to-sequence tasks such as translation. Auto-regressive possibly and adjustment of the GUI.
models predict the next single token (sub-word) by max- In the following sections, we demonstrate how to
elimimizing the log-likelihood given all previous words and inate the need for a product-specific GUI by utilizing an
the model parameters [4]. LLM to engage in dialogue with the user.</p>
          <p>The size and quality of the pre-training corpus have
a strong impact on performance [5]. LLMs are trained
on publicly available data and excel in general language 3. Related Work
tasks. Highly specialized tasks require expert
knowledge that is often not included in the training data, and
therefore LLMs may not be able to generate accurate
output. Task-specific knowledge can be introduced to a
general-purpose LLM through domain customization by
employing techniques like prompting and fine-tuning [6].</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Various approaches to improve the reliability, the per</title>
          <p>formance in domain-specific tasks, and the reasoning
abilities of LLMs are described in literature.</p>
          <p>Few-shot prompting efectively introduces
domainspecific knowledge and improves the task-specific
performance of LLMs by adding a small set of example
interactions (input and expected output) to the prompt [10].
2.2. Constraint-based Product Chain-of-thought prompting was shown to improve the
Configuration reasoning abilities of LLMs especially in more complex
tasks by providing exemplary intermediate reasoning
Product configuration involves selecting and assembling steps [11].
various components and options to meet customer re- Grammar prompting is used when a specific output
quirements and constraints. Its complexity arises from format is expected. Wang et al. describe how a minimal
the vast number of possible combinations and the need specialized grammar is obtained in a grammar
specialto satisfy all technical restrictions and customer prefer- ization process by selecting a specialized grammar as a
ences. To handle this complexity, powerful technolo- subset of the full grammar using an LLM and
minimizgies have been developed and established in the last ing it by parsing the output and forming the union of
decades. Constraint-based systems shall be highlighted used rules. In their approach, constrained decoding then
here, which allow to represent the product line and its validates the output syntax [12]. Similarly, Poesia et al.
technical restrictions and requirements in a clean, logical presented the Synchromesh framework: Using a few-shot
way, thereby ensuring that only valid configurations are prompting technique, semantically similar examples are
generated. The core of such systems lies in the ability selected from a larger pool for a given natural language
to handle complex and combinatorial search spaces efi- prompt via a similarity metric named Target Similarity
ciently through the use of advanced solving algorithms, Tuning. Constraints are enforced through Constrained
such as backtracking, forward checking, and constraint Semantic Decoding to verify syntax validity, scoping, or
propagation. This facilitates the eficient generation of type checks. During the token-by-token construction of
feasible solutions while pruning invalid combinations. the LLM output, a Completion Engine provides all valid</p>
          <p>An important subdomain of configuration problems tokens that can further extend a partial program towards
are feature models for the representation of product a full correct program [13].
lines [7]. Constraint-based techniques are especially well- Neuro-symbolic approaches focus on combining the
suited for such feature models, because of the simple strengths of neural networks and symbolic reasoners.
language and the mainly Boolean type of the variables. Pan et al. introduced the Logic-LM framework which</p>
          <p>MiniZinc is a constraint language that can be used to achieves a performance improvement of 18% on logical
represent configuration problems [ 8]. Several eficient reasoning datasets over chain-of-though prompting. The
solvers can process this language and can therefore be framework translates the natural-language input into
used as the backend of a configurator. symbolic formulations and utilizes a symbolic reasoner</p>
          <p>A product configurator is almost always an interactive to obtain the answer [14].
system [9]. A graphical user-interface (GUI) allows to This paper builds upon our previous work [15] that
enter the user requirements, which are passed on as input
studied the reliable generation of formal specifications
with LLMs using algorithmic post-processing. We extend
the approach towards product configuration by applying
post-processing to reliably integrate a constraint solver.</p>
          <p>In addition to previously described guaranteed
syntactically valid output, this extension enables arbitrary
semantic constraints.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Configuration Copilot</title>
      <sec id="sec-4-1">
        <title>This section presents the technical details of the configuration copilot that combines LLMs with constraint-based configuration.</title>
        <sec id="sec-4-1-1">
          <title>4.1. Architecture</title>
        </sec>
        <sec id="sec-4-1-2">
          <title>4.2. Formalizer</title>
          <p>The input to the Formalizer is a natural-language
description of arbitrary product requirements provided by
a non-expert. Utilizing the NLP capabilities of LLMs, the
formalization can be viewed as a sequence-to-sequence Few-shot prompting has been shown to efectively
extranslation task from natural language to a formal speci- tend the capabilities of LLMs with domain knowledge
ifcation. The LLM is tasked with natural language under- while requiring significantly less training data than
finestanding and the identification of corresponding parame- tuning [10]. Knowledge of the product line is
incorpoters or components of the product (line), but is specifically rated through a system prompt describing the product
not tasked with reasoning (e.g., constraint satisfaction). line with its parameters and components. A small set
While pre-trained LLMs achieve a strong performance on of examples is appended as pairs of natural-language
ingeneral tasks, they do not have knowledge of the specific puts and expected outputs to provide the LLM with more
product (line) to configure as corresponding data is not context and guide it towards the expected behavior.
included in their training corpus [5, 6]. Additionally, the Rather than generating output directly in a specific
probabilistic nature of the token-by-token output con- constraint language, an intermediary JSON-based
lanstruction of LLMs does not provide any guarantees in guage is used, which can then be easily transpiled. The
the correct generation of valid constraints [4]. transpiler parses the JSON constraint representation and</p>
          <p>
            Our framework for reliable code generation addresses maps its elements to corresponding constructs of the
domain-customization and reliable output generation specific constraint language following predefined rules.
through few-shot prompting and algorithmic post- As JSON is widely used, pre-trained LLMs have more
processing [15]. often encountered JSON than less common constraint
languages. Therefore, the generation of an intermediary a configuration. The product line as well as the user
JSON output is closer to the LLMs capabilities. Addition- constraints are modelled in the MiniZinc constraint
lanally, an intermediary language gives more control over guage [8]. The solver returns the full product
configurathe expected output as available language constructs can tion as a list of variable assignments which serves as an
be constrained and tailored to the specific task. It also input to the Interpreter. In this context, we consider the
decouples the Formalizer from the Configuration Engine constraint solver a given technology that will neither be
by enabling interchangeability of the concrete constraint further described nor evaluated.
language. To generate a valid JSON for the Formalizer,
several state-of-the-art LLMs are evaluated and bench- 4.4. Interpreter
marked. Specialized code LLMs that are pre-trained on
the translation of natural language to code in a variety of The Interpreter is an LLM module that explains the
prodprogramming languages are believed to be more suitable uct configuration found by the Configuration Engine.
for the generation of structured JSON output. In our eval- The goal is to provide the user with a less technical
sumuation in Section 5, we selected four open-access LLMs: mary that is understandable for non-experts.
Two code LLMs (CodeLLama [16] and Codestral [17]), Structured few-shot prompting [10] is suficient for
and two general-purpose instruction-tuned LLMs (Meta this use-case as LLMs generally perform well in the
transLlama 3 [18] and Mistral [
            <xref ref-type="bibr" rid="ref3">19</xref>
            ]). lation from a formal specification to a natural-language
          </p>
          <p>Algorithmic post-processing guarantees the correct summary as all facts are directly present in the prompt.
generation of the JSON-based intermediary language and The context given to the LLM consists of three aspects:
is depicted in Figure 2. As the auto-regressive Trans- The product line definition, instructions, and examples.
former model generates its output step-by-step as tokens, The LLM is prompted to evaluate which properties and
the post-processor engages into every generation step: components are most important to be included in the
For each step, the LLM generates a list of candidates for summary. This is achieved by adding importance hints
the next token based on the prompt and the generated to the product line definition, and by appending the
origoutput so far. Sorted by priority as evaluated by the inal user input. Properties and components mentioned
LLM, the post-processor determines whether the token directly in the user input are given more importance and
candidate represents a valid continuation of the partial are more likely to be included in the summary. The result
output sequence (partial intermediary JSON). The valid is a more natural context-aware explanation of the most
token candidate with the highest priority is then selected, relevant aspects in the product configuration.
handed back to the LLM, and added to the partial JSON,
extending it one step further. A completeness checker 5. Evaluation
evaluates after every step to determine, if the JSON is
complete [15]. The presented configuration copilot is evaluated on two</p>
          <p>The JSON-based intermediary language is formally use-cases: The conceptually simpler task of configuring
defined by a JSON schema specification and the post- a feature model, and the configuration of a metro Wagon.
processor is therefore a specialized JSON validator that
can strictly validate any partial JSON against the schema.</p>
          <p>
            This implementation is based on deterministic finite au- 5.1. Feature Model (GoPhone)
tomata (DFA). Each generic JSON language element (ob- The first use-case for the evaluation of the presented
ject, list, string, number, etc.) is represented by a DFA, copilot is the configuration of a feature model. An
uninkeeping track of the current state. The token generated formed user shall be supported in the configuration of
by the LLM is broken down to single-character inputs the GoPhone from the SPLOT project [
            <xref ref-type="bibr" rid="ref4">20</xref>
            ].
for the JSON validator. Depending on the schema and The GoPhone is a feature model comprised of 77
feathe current state, only a set of characters is accepted. If tures with some being mandatory, optional, dependent
a character is rejected, the current token is considered on other features, or mutually exclusive. For example,
invalid, and the validator state is rolled back to the last the feature call is mandatory for the GoPhone, the
feavalid token. State changes are triggered by characters ture accept_incoming_call is mandatory for call, but
until the final state is reached. When the DFA reaches show_missed_calls and show_received_calls are
opits final state, the generated valid JSON is complete [ 15]. tional.
          </p>
        </sec>
        <sec id="sec-4-1-3">
          <title>4.3. Configuration Engine</title>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Given the user constraints combined with the complete product line definition, the Configuration Engine evaluates whether the constraints are satisfiable and returns</title>
      </sec>
      <sec id="sec-4-3">
        <title>Feature assignments are Boolean, either the feature is included in the product configuration ( true) or the feature is not included (false). The product line definition is a MiniZinc program that was directly derived from the</title>
        <p>feature model. Each feature is a Boolean variable. Con- Together with the MiniZinc program (product line
straints limit the combination of features and therefore definition), the Configuration Engine evaluates the
conlimit possible product configurations. straints and returns a full product configuration of</p>
        <p>
          A non-expert user starts by describing their require- the GoPhone for the specific user requirements as a
ments for the phone in natural language: list of Boolean feature assignments. In this work, the
Gecode [
          <xref ref-type="bibr" rid="ref5">21</xref>
          ] solver was used without further
configuI need a basic phone to call people and browse ration or optimization. The Interpreter converts this
tkheeepwetbrabcuktoIf dmoyn'atpppolianytmegnatmse.s. I also want to configuration back to natural language and returns it to
the user. An example for such an output is (the technical
specification is shortened for brevity):
{
"features": [
{
},
{
},
{
},
{
}
"name": "make_call",
"value": true
"name": "browsing",
"value": true
"name": "game",
"value": false
"name": "calendar_entry",
"value": true
        </p>
        <p>Your GoPhone can manage ringing tones, messages,
and browse the web. It can also manage calls,
read multimedia, and display photos. It has a
calendar entry feature and an address book
processing system. However, it does not play
games, organize tasks, or have currency
conversion features.</p>
        <p>Here is the full technical configuration:
GoPhone = true;
manage_ringing_tones = true;
[...]
browse = true;
[...]
game = false;
play_games = false;
install_games = false;
[...]
} The crucial and potentially failing component of the
presented architecture is the Formalizer: the probabilistic</p>
        <p>This list of solver-independent constraints is then tran- nature of the underlying LLMs does not provide strict
spiled to MiniZinc constraints: guarantees. Especially the translation of the user’s
requirements to the feature assignments is subject to
unconstraint make_call = true; certainty. A formal evaluation of the Interpreter is not
constraint browsing = true; done because the correctness requirements for the
configconstraint game = false; uration summary are less strong and LLMs are generally
constraint calendar_entry = true; known to perform well on simple summarization tasks
when the facts are directly provided. It is also unsuitable
to define a single reference solution as a large variety Table 1
of summaries (with various feature assignments being Evaluation Results for the GoPhone Formalization
explained or not explained) could be considered correct. S = Similarity score
Ultimately, users needs to decide whether the summary F1 = F1 score
was helpful or not. Model [Size/Quantization] S F1</p>
        <p>The Formalizer was evaluated on a custom dataset of CodeLlama 34B/Q4 [Link] 0.65 0.74
30 test cases. 15 test cases create a new configuration Codestral 22B/Q4 [Link] 0.79 0.86
from scratch, and 15 test cases evaluate a re-configuration Meta Llama 3 8B/Q8 [Link] 0.46 0.58
where a given configuration is modified. Each test case Mistral 7B/Q8 [Link] 0.69 0.79
consists of natural-language input mentioning between
two and six feature requirements in the text (and up to The overall similarity between the expected and actual
30 given feature assignments for modification test cases),
and the expected feature assignments in JSON. Using the output  as the weighted average of   and   is:
natural-language input, the Formalizer generates feature   ⋅ | ∪  | ̂ +   ⋅ | ∪  | ̂
assignments in JSON. This output is compared to the ex-  =
pected output. The comparison is conceptually challeng- | ∪  | ̂ + | ∪  | ̂
ing due to the intrinsic ambiguity of natural language. In The result is a number between 0 and 1, with 0
indimany cases, one could argue for multiple options of fea- cating no similarity, and 1 indicating a perfect match. In
ture assignments to be considered a correct translation. this metric, the identification of features in the natural
In this evaluation, we hand-crafted the dataset to be less language as well as the Boolean assignment are
considambiguous. However, the features of the GoPhone are ered.
in themselves sometimes not obviously distinguishable, Similarly, the precision  , recall  and F1 score  1 were
and multiple features may be equally suitable. For exam- calculated:
ple, the feature browsing is an optional sub-feature of
the more general parent feature browse. This ambiguity  = | ∩  | ̂ + | ∩  | ̂
was addressed by encoding very similar features to the | | ̂ + |  | ̂
same representation. Therefore, all defined synonymous
features are considered a correct feature assignment for  = | ∩  | ̂ + | ∩  | ̂
a requirement. However, the feature assignment was | | + | |
rneoatsloinmiintgedretqouleiraefmfeeanttusrteos tbheecFauorsmedaloiiznegr. sCoownsoiudledr atdhde  1 = 2 ⋅   +⋅  
leaf features play_games and install_games, and the We selected four open-access LLMs from HuggingFace
parent feature game. If a user only mentions games in to be evaluated in the context of the configuration copilot,
their descriptions, the more abstract feature game shall two code models, and two general-purpose models.
Tabe assigned. Otherwise, the LLM would have to reason ble 1 summarizes the evaluation results for the GoPhone
about a proper assignment of leaf features, deviating use-case per LLM. Codestral 22B/Q4, a state-of-the-art
from the most direct translation from natural language code model, performed best. However, Mistral 7B/Q8
to a feature assignment. The reasoning regarding further outperformed the larger code model CodeLLama 34B/Q4
(sub-)feature assignments shall be done by the Configu- against our expectations. This shows that the
perforration Engine. mance of LLMs is use-case specific and must be evaluated.</p>
        <p>
          A similarity metric based on the Jaccard distance be- We found that the performance degrades as instances
between sets [
          <xref ref-type="bibr" rid="ref6">22</xref>
          ] was used to compare each pair of expected come more complex. Remedies for this observation are
and actual output: Let  and  be the sets of feature the use of larger models, tuning the technical approach,
names in the expected output where the feature value or future improvements of LLMs themselves.
Considis    and   , respectively. Similarly, let  ̂ and  ̂ be ering the remaining ambiguity of natural language, the
the sets of feature names in the actual output where the results indicate reasonable performance in this use-case
feature value is    and   , respectively. The Jaccard as the majority of feature requirements was formalized
similarities are: correctly.
        </p>
        <p>For the    sets:
For the  
sets:
  =
  =
| ∩  | ̂
| ∪  | ̂
| ∩  | ̂
| ∪  | ̂</p>
        <sec id="sec-4-3-1">
          <title>5.2. Metro Wagon</title>
        </sec>
      </sec>
      <sec id="sec-4-4">
        <title>The second use-case for the evaluation is a metro Wagon configuration problem (see [ 23]) that uses not only Boolean but also numeric variables and arrays, where a configurable product has components that can occur</title>
        <p>length_mm: 10000...20000
nr_passengers: 50..200
nr_seats: 0..200
standing_room: 0..200
nr_seats + standing_room = nr_passengers
nr_seats + standing_room/3
≤ 4*length_mm/1000
nr_seats = count(Seat)
standing_room&gt;0 → count(Handrail)=1
all-equal-type()
all-equal-color()
maximize nr_passengers/length_mm
0..1</p>
        <p>Handrail
type: {standard,
premium}
0..80</p>
        <p>Seat
type: {standard,
premium, special}
color: {blue, red, white}
type=special → color=red</p>
      </sec>
      <sec id="sec-4-5">
        <title>Together with the MiniZinc program (product line</title>
        <p>
          definition), the Configuration Engine evaluates the
conmultiple times (similar to generative constraint satisfac- straints and returns a full product configuration of the
tion [
          <xref ref-type="bibr" rid="ref8">24</xref>
          ] or cardinality-based feature modelling [
          <xref ref-type="bibr" rid="ref9">25</xref>
          ]). metro Wagon for the specific user requirements as a list
        </p>
        <p>A metro train wagon has as configurable attributes of value assignments to the configurable parameters. The
the size (length in millimetres: 10000..20000) and the Interpreter converts this configuration back to natural
expected load (number of passengers: 50..200) which can language and returns it to the user:
be realized as seats or standing room. As components Your metro Wagon is 20 meters long, has space for
we consider only seats (max. 4 per meter of length) and 160 passengers with 40 red standard seats and a
handrails, and their number is configurable. standard handrail. There is also standing room</p>
        <p>There is at most one handrail in a wagon (mandatory for an additional 120 people.
if there is standing room) and it has a configurable type:
“standard” or “premium”. Here is the full technical configuration:</p>
        <p>A single seat consumes standing room for 3 persons
and has as configurable attributes the type (“standard”, length_mm = 20000;
“premium”, “special”) and the color (“blue”, “red”, “white”). nr_passengers = 160;
The type is constrained such that standard is not allowed sntra_nsdeiantgs_=ro4o0m;= 120;
to be mixed with premium (for seats and handrails). The nr_handrails = 1;
color of all seats must be the same, except for special handrail_type = standard;
seats which have to be “red”. seat_color = [red, red, red, ...];</p>
        <p>Figure 3 shows a UML class diagram for this sample seat_type = [standard, standard, standard, ...];
specification, including pseudo code for all constraints.</p>
        <p>A non-expert user starts by describing their
requirements for the metro Wagon in natural language:
The wagon should accommodate more than 120 people
with room for 40 to sit. Seats should be red.</p>
      </sec>
      <sec id="sec-4-6">
        <title>This natural-language description is then formalized</title>
        <p>to the intermediary JSON language:</p>
        <p>The Formalizer for the metro Wagon use-case was
evaluated, like the GoPhone Formalizer, on a diverse
set of 30 test cases (pairs of input and expected output)
with 15 creating a new configuration and 15 modifying
a given configuration (re-configuration). To evaluate
the similarity in this use-case, the previously described
similarity metric based on the Jaccard distance between
sets for Boolean feature assignments was extended to
the more general use-case. This extension is necessary
to enable the evaluation of the value assignment for the
• for array values: length and positional item equal- Formalizer, was evaluated, a formal evaluation of the
In</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6. Conclusion</title>
      <p>This paper presented a configuration copilot that enables
non-expert users to configure a product in natural
language. The cooperative neuro-symbolic approach
combines an LLM with a constraint solver to reliably support
a product configuration. An early evaluation on the two
use-cases of configuring the GoPhone feature model and
a metro Wagon indicated practical feasibility. We believe
that a configuration copilot is a valuable extension to
extended variable types (i.e., strings, numbers, arrays).</p>
      <sec id="sec-5-1">
        <title>GUI-based product configurators. For a productive im</title>
        <p>While the Jaccard distance remained the basis for the
plementation, limitations and future work mentioned in
according to the value similarities as well. The type- A limitation of our work is the size of the use-cases.
Comsimilarity metric, a type-specific value metric was applied
to each configuration parameter that is present in both,
the expected and the actual output. In addition to the
parameters being present, the total similarity is adjusted
specific metric considers:
• for numeric values: operator (’=’, ’&gt;’, ’&lt;’, etc.) and
value distance relative to the parameter-specific
domain (value range)
ity
• for string-enumerated values: exact value match</p>
        <p>Let  be the set of configuration parameter names in
the expected output and let  ̂ be the set of
configuration parameter names in the actual output. The Jaccard
similarity   is:
  =
| ∩ | ̂
| ∪ | ̂</p>
      </sec>
      <sec id="sec-5-2">
        <title>Let  be a matching parameter that is in both, the ex</title>
        <p>pected and the actual output, and let   () be the
typespecific value similarity (between 0 and 1) of  between
the expected and the actual output. The value-adjusted</p>
      </sec>
      <sec id="sec-5-3">
        <title>Jaccard similarity  is then:</title>
        <p>=
∑ ∈ ∩  ̂  ()</p>
        <p>| ∪ | ̂</p>
        <p>The evaluation of the F1 score is omitted because it
does not provide any additional value, as it appears to
correlate strongly with the already rather strict similarity
score  . Table 2 summarizes the evaluation results for
the metro Wagon use-case per LLM. Codestral 22B/Q4
performed best again with a similar score. However, the
other three models consistently improved their score
compared to the GoPhone use-case. While the metro
usecase in itself is more complex, the domain size (amount
of named parameters) is lower, which may be the reason
for the higher performance. Overall, the results again
indicate a reasonable performance for the metro Wagon
use-case.</p>
      </sec>
      <sec id="sec-5-4">
        <title>Sections 6.1 and 6.2 should be addressed.</title>
        <sec id="sec-5-4-1">
          <title>6.1. Limitations</title>
          <p>pared to real-world scenarios, the evaluated GoPhone
feature model and metro Wagon are smaller and less
complex. Additionally, the evaluation was done on a limited
manually created dataset with 30 instances per use-case.
While the most critical aspect of the architecture, the
terpreter and the full configuration pipeline were omitted
for the reason that a user study is required to evaluate
these aspects. This paper demonstrates that creating a
productive configuration copilot is feasible but does not
study the extent to which value is provided to real users
in a real-world scenario.</p>
        </sec>
        <sec id="sec-5-4-2">
          <title>6.2. Future Work</title>
          <p>To address the limitations of this paper, the configuration
copilot shall be evaluated on more complex use-case from
practice in a user study. The configuration copilot itself
shall be extended: When a configuration as specified
by the user is unsatisfiable, the configuration copilot
shall suggest alternatives instead of reverting to the last
satisfiable configuration. Additionally, soft constraints
in the form of ’If possible, I would like to ...’ shall be
introduced.
commerce environment: The impact of product
configurator interaction design on consumer
personalized customization experience, Sustainability
14 (2022). URL: https://www.mdpi.com/2071-1050/
[3] E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, D. Amodei, Language models are few-shot
Y. Zhou, S. Savarese, C. Xiong, Codegen: An open learners, in: H. Larochelle, M. Ranzato, R. Hadsell,
large language model for code with multi-turn pro- M. Balcan, H. Lin (Eds.), Advances in Neural
Inforgram synthesis, in: The Eleventh International Con- mation Processing Systems, volume 33, Curran
ference on Learning Representations, ICLR 2023, Associates, Inc., 2020, pp. 1877–1901. URL: https:
Kigali, Rwanda, May 1-5, 2023, OpenReview.net, //proceedings.neurips.cc/paper_files/paper/2020/
2023. file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
[4] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, [11] J. Wei, X. Wang, D. Schuurmans, M. Bosma,
L. Jones, A. N. Gomez, L. u. Kaiser, I. Polosukhin, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, D. Zhou,
ChainAttention is all you need, in: I. Guyon, U. V. of-thought prompting elicits reasoning in large
lanLuxburg, S. Bengio, H. Wallach, R. Fergus, guage models, in: Proceedings of the 36th
InternaS. Vishwanathan, R. Garnett (Eds.), Advances tional Conference on Neural Information
Processin Neural Information Processing Systems, vol- ing Systems, NIPS ’22, Curran Associates Inc., Red
ume 30, Curran Associates, Inc., 2017. URL: https: Hook, NY, USA, 2024.
//proceedings.neurips.cc/paper_files/paper/2017/ [12] B. Wang, Z. Wang, X. Wang, Y. Cao, R. A. Saurous,
file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf. Y. Kim, Grammar prompting for domain-specific
[5] B. Min, H. Ross, E. Sulem, A. P. B. Veyseh, T. H. language generation with large language models,
Nguyen, O. Sainz, E. Agirre, I. Heintz, D. Roth, Re- in: A. Oh, T. Neumann, A. Globerson, K. Saenko,
cent advances in natural language processing via M. Hardt, S. Levine (Eds.), Advances in Neural
Inforlarge pre-trained language models: A survey, ACM mation Processing Systems, volume 36, Curran
AsComputing Surveys (2023). sociates, Inc., 2023, pp. 65030–65055. URL: https://
[6] C. Ling, X. Zhao, J. Lu, C. Deng, C. Zheng, proceedings.neurips.cc/paper_files/paper/2023/file/
J. Wang, T. Chowdhury, Y. Li, H. Cui, X. Zhang, cd40d0d65bfebb894ccc9ea822b47fa8-Paper-Conference.
T. Zhao, A. Panalkar, W. Cheng, H. Wang, Y. Liu, pdf.</p>
          <p>Z. Chen, H. Chen, C. White, Q. Gu, C. Yang, [13] G. Poesia, A. Polozov, V. Le, A. Tiwari, G. Soares,
L. Zhao, Beyond one-model-fits-all: A survey C. Meek, S. Gulwani, Synchromesh: Reliable code
of domain specialization for large language mod- generation from pre-trained language models, in:
els, CoRR abs/2305.18703 (2023). URL: https:// The Tenth International Conference on Learning
doi.org/10.48550/arXiv.2305.18703. doi:10.48550/ Representations, ICLR 2022, Virtual Event, April
ARXIV.2305.18703. arXiv:2305.18703. 25-29, 2022, OpenReview.net, 2022.
[7] D. Benavides, A. Felfernig, J. A. Galindo, F. Rein- [14] L. Pan, A. Albalak, X. Wang, W. Wang,
Logicfrank, Automated analysis in feature modelling and LM: Empowering large language models with
symproduct configuration, in: Safe and Secure Software bolic solvers for faithful logical reasoning, in:
Reuse: 13th International Conference on Software H. Bouamor, J. Pino, K. Bali (Eds.), Findings of the
Reuse, ICSR 2013, Pisa, June 18-20. Proceedings 13, Association for Computational Linguistics: EMNLP
Springer, 2013, pp. 160–175. 2023, Association for Computational Linguistics,
[8] N. Nethercote, P. J. Stuckey, R. Becket, S. Brand, G. J. Singapore, 2023, pp. 3806–3824. URL: https://
Duck, G. Tack, MiniZinc: Towards a standard CP aclanthology.org/2023.findings-emnlp.248. doi:10.
modelling language, in: CP, volume 4741 of LNCS, 18653/v1/2023.findings-emnlp.248.</p>
          <p>Springer, 2007, pp. 529–543. [15] P. Kogler, A. Falkner, S. Sperl, Reliable
genera[9] A. A. Falkner, A. Haselböck, G. Krames, G. Schenner, tion of formal specifications using large language
R. Taupe, Constraint solver requirements for inter- models, in: SE 2024 - Companion, Gesellschaft für
active configuration, in: L. Hotz, M. Aldanondo, Informatik e.V., 2024, pp. 141–153. doi:10.18420/
T. Krebs (Eds.), Proceedings of the 21st Config- sw2024-ws_10.
uration Workshop, Hamburg, Germany, Septem- [16] B. Rozière, J. Gehring, F. Gloeckle, S. Sootla,
ber 19-20, 2019, volume 2467 of CEUR Workshop I. Gat, E. Tan, Y. Adi, J. Liu, T. Remez, J. Rapin,
Proceedings, CEUR-WS.org, 2019, pp. 65–72. URL: A. Kozhevnikov, I. Evtimov, J. Bitton, M. Bhatt,
http://ceur-ws.org/Vol-2467/paper-12.pdf. C. C. Ferrer, A. Grattafiori, W. Xiong, A.
Défos[10] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. sez, J. Copet, F. Azhar, H. Touvron, L. Martin,
Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, N. Usunier, T. Scialom, G. Synnaeve, M. Ai, Code
G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, llama: Open foundation models for code, 2023. URL:
G. Krueger, T. Henighan, R. Child, A. Ramesh, https://github.com/facebookresearch/codellama.
D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, [17] MistralAI, Codestral introduction (2024). URL:
E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, https://mistral.ai/news/codestral/.
C. Berner, S. McCandlish, A. Radford, I. Sutskever, [18] AI@Meta, Llama 3 model card (2024). URL:</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>L. Zhang,</surname>
          </string-name>
          <article-title>Product configuration: A review of the state-of-the-art and future research</article-title>
          ,
          <source>International Journal of Production Research</source>
          <volume>52</volume>
          (
          <year>2014</year>
          )
          <fpage>6381</fpage>
          -
          <lpage>6398</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Yi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yu</surname>
          </string-name>
          , Creating a sustainable ehttps://github.com/meta-llama/llama3/blob/main/ MODEL_CARD.md.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A. Q.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sablayrolles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mensch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bamford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Chaplot</surname>
          </string-name>
          , D. de las Casas,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bressand</surname>
          </string-name>
          , G. Lengyel,
          <string-name>
            <given-names>G.</given-names>
            <surname>Lample</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Saulnier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. R.</given-names>
            <surname>Lavaud</surname>
          </string-name>
          , M.
          <article-title>-</article-title>
          <string-name>
            <surname>A. Lachaux</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Stock</surname>
            ,
            <given-names>T. L.</given-names>
          </string-name>
          <string-name>
            <surname>Scao</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Lavril</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Lacroix</surname>
            ,
            <given-names>W. E.</given-names>
          </string-name>
          <string-name>
            <surname>Sayed</surname>
          </string-name>
          , Mistral 7b,
          <year>2023</year>
          . arXiv:
          <volume>2310</volume>
          .
          <fpage>06825</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mendonca</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Branco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cowan</surname>
          </string-name>
          , S.p.l.o.t.
          <article-title>- software product lines online tools</article-title>
          ,
          <source>In Companion to the 24th ACM SIGPLAN International Conference on Object-Oriented Programming Systems, Languages, and Applications</source>
          , OOPSLA (
          <year>2009</year>
          )
          <fpage>761</fpage>
          -
          <lpage>762</lpage>
          . doi:
          <volume>10</volume>
          .1145/1639950.1640002.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Gecode</surname>
            <given-names>Team</given-names>
          </string-name>
          ,
          <article-title>Gecode: Generic constraint development environment</article-title>
          ,
          <year>2006</year>
          . Available from http://www.gecode.org.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>M.</given-names>
            <surname>LEVANDOWSKY</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. WINTER</surname>
          </string-name>
          ,
          <article-title>Distance between sets</article-title>
          ,
          <source>Nature</source>
          <volume>234</volume>
          (
          <year>1971</year>
          )
          <fpage>34</fpage>
          -
          <lpage>35</lpage>
          . URL: https:// doi.org/10.1038/234034a0. doi:
          <volume>10</volume>
          .1038/234034a0.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>A.</given-names>
            <surname>Falkner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Haselböck</surname>
          </string-name>
          , G. Krames, G. Schenner,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schreiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Comploi-Taupe</surname>
          </string-name>
          ,
          <article-title>Solver requirements for interactive configuration</article-title>
          ,
          <source>JOURNAL OF UNIVERSAL COMPUTER SCIENCE 26</source>
          (
          <year>2020</year>
          )
          <fpage>343</fpage>
          -.
          <source>doi:10</source>
          .3897/jucs.
          <year>2020</year>
          .
          <volume>019</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>G.</given-names>
            <surname>Fleischanderl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Haselböck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schreiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stumptner</surname>
          </string-name>
          ,
          <article-title>Configuring large systems using generative constraint satisfaction</article-title>
          ,
          <source>IEEE Intelligent Systems</source>
          <volume>13</volume>
          (
          <year>1998</year>
          )
          <fpage>59</fpage>
          -
          <lpage>68</lpage>
          . URL: https://doi.org/10.1109/5254.708434. doi:
          <volume>10</volume>
          .1109/ 5254.708434.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>K.</given-names>
            <surname>Czarnecki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Helsen</surname>
          </string-name>
          , U. W. Eisenecker,
          <article-title>Formalizing cardinality-based feature models and their specialization</article-title>
          ,
          <source>Software Process: Improvement and Practice</source>
          <volume>10</volume>
          (
          <year>2005</year>
          )
          <fpage>7</fpage>
          -
          <lpage>29</lpage>
          . URL: https://doi.org/ 10.1002/spip.213. doi:
          <volume>10</volume>
          .1002/spip.213.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>