<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam,
G. Sastry, A. Askell, et al., Language models are few-shot learners, Advances in Neural Information
Processing Systems</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>ATLAS: Benchmarking and Adapting LLMs for Global Trade via Harmonized Tarif Code Classification ⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pritish Yuvraj</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Siva Devarakonda</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2026</year>
      </pub-date>
      <volume>33</volume>
      <issue>2020</issue>
      <fpage>27730</fpage>
      <lpage>27744</lpage>
      <abstract>
        <p>Accurate classification under the Harmonized Tarif Schedule (HTS) is a critical yet underexplored problem in global trade compliance, where errors can delay shipments and disrupt supply chains. We present Atlas, the first benchmark and fine-tuned large language model for HTS code prediction, constructed from the U.S. Customs Rulings Online Search System (CROSS). The benchmark includes 18,731 legally grounded rulings spanning 2,992 unique codes, reformatted into reasoning-oriented prompts. Our fine-tuned Atlas model (LLaMA-3.3-70B) achieves 40% accuracy at the full 10-digit level and 57.5% at the 6-digit level-improvements of +15 and +27.5 points over strong baselines-while being approximately 5× cheaper to deploy. These results establish HTS classification as a rigorous benchmark for hierarchical reasoning, cost-eficient adaptation, and alignment in domain-specialized large language models. The dataset and model are publicly released to encourage further research on structured reasoning for real-world compliance tasks.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;large language models</kwd>
        <kwd>hierarchical reasoning</kwd>
        <kwd>benchmark</kwd>
        <kwd>domain adaptation</kwd>
        <kwd>fine-tuning</kwd>
        <kwd>trade compliance</kwd>
        <kwd>tarif classification</kwd>
        <kwd>HTS code</kwd>
        <kwd>LLaMA</kwd>
        <kwd>structured prediction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>1.1. Contributions</title>
        <sec id="sec-1-1-1">
          <title>We focus on the high-value semiconductor domain and present:</title>
          <p>• The first open-source benchmark for HTS classification [
Online Search System (CROSS);</p>
        </sec>
        <sec id="sec-1-1-2">
          <title>5], derived from the U.S. Customs Rulings</title>
          <p>• A comprehensive evaluation of leading proprietary and open-source models, including
GPT-5</p>
          <p>Thinking, Gemini-2.5-Pro-Thinking, LLaMA-3.3-70B, DeepSeek-R1, and GPT-OSS-120B;
• The fine-tuned Atlas model [ 6] (LLaMA-3.3-70B), achieving 40% 10-digit and 57.5% 6-digit
accuracy—substantially outperforming baselines—while being up to 8× cheaper and fully
selfhostable for privacy-sensitive deployments.</p>
          <p>Together, these contributions position tarif code classification as a new benchmark for evaluating
reasoning and adaptation in large language models. Traditional approaches to product classification
typically rely on hand-crafted features (e.g., HS keyword rules, tarif-engine heuristics) or shallow
models such as random forests and gradient-boosted trees trained on bag-of-words representations.
While such models can work well for constrained taxonomies, they struggle to scale to thousands of
ifne-grained, legally nuanced classes and ofer limited support for explanation, counterfactual analysis,
or rapid adaptation to regulatory change. In contrast, large language models can ingest raw ruling text,
reason over subtle distinctions (e.g., “partially fabricated” vs. “finished” wafers), and produce both a
code and an accompanying rationale. A systematic comparison with non-LLM baselines is an important
direction for future work, but in this paper we focus on establishing a strong LLM-based baseline and a
reusable benchmark.</p>
          <p>Relevance to Knowledge Graphs and Agentic Systems. Although this work focuses on supervised
ifne-tuning, HTS classification naturally interfaces with agentic systems and knowledge graphs. In
practical deployments, a tarif-classification agent must (i) retrieve and ground its predictions in
structured regulatory corpora (e.g., HTS knowledge graphs linking codes to legal notes and duty rates),
(ii) orchestrate multi-step reasoning workflows (e.g., querying prior rulings, edge-case escalation, and
audit trails), and (iii) interact with downstream customs-compliance pipelines. We position Atlas as a
core reasoning component in such agentic systems: our benchmark and model provide a standardized,
hierarchical decision layer that can be plugged into knowledge-graph–augmented retrieval and
toolusing agents for end-to-end trade compliance.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Dataset</title>
      <p>Our main contribution is the first large-scale dataset for Harmonized Tarif Schedule (HTS) classification,
derived from the U.S. Customs Rulings Online Search System (CROSS) [7]. CROSS contains legally
binding rulings by U.S. Customs and Border Protection (CBP) specifying the correct 10-digit HTS code
for products. These rulings are authoritative yet dispersed across thousands of HTML pages, previously
inaccessible for ML research.</p>
      <p>Example instance. A typical CROSS ruling in our dataset describes, for example, a shipment of
“12-inch silicon wafers that have undergone photolithographic patterning but are not yet diced or
packaged.” The corresponding HTS US code is 8542.90.0100, which encodes (roughly) “electronic
integrated circuits; other; wafers; unmounted.” Our processed instance contains (i) a condensed product
description, (ii) a reasoning-style explanation derived from the ruling (e.g., why certain headings are
excluded), and (iii) the final 10-digit code.</p>
      <sec id="sec-2-1">
        <title>2.1. Collection and Scope</title>
        <p>We built an automated agent [8, 9, 10] to scrape CROSS and align each ruling with its oficial 10-digit
HTS code from [1]. Focusing on semiconductor and manufacturing chapters, we obtained 18,731 rulings
covering 2,992 unique codes. Frequent rulings highlight ambiguous or high-demand categories, while
absent codes suggest stable ones. Table 1 lists representative chapters (the complete distribution is
provided in Appendix A, Table 5).</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Transformation to Model-Ready Format</title>
        <p>Raw rulings are verbose legal letters. We used GPT-4o-mini [11] for information extraction, converting
each into a concise instruction–response pair containing (a) a product description, (b) reasoning trace,
and (c) final HTS code. The complete prompt template is provided in Appendix B. This structure
enforces reasoning-based prediction, aligning with chain-of-thought research [12].</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Splits and Availability</title>
        <p>We reserved 200 samples each for validation and testing, with the remaining 18,254 for training
(Table 2). To ensure fair and representative evaluation despite the small test size, the 200 test samples
were stratified across high-variance HTS chapters (e.g., 84, 85, and 90) to reflect the diversity and
ambiguity observed in real-world tarif rulings. The dataset is publicly available on Hugging Face [5].</p>
        <p>We chose relatively small validation and test splits (200 examples each) for two reasons. First, manual
inspection and error analysis at the 10-digit level are time-consuming because each prediction must be
checked against lengthy legal notes and chapter-specific carve-outs. Second, our goal was to maximize
the training signal for Atlas while still preserving a stratified and diverse test set that covers
highvariance chapters (e.g., 84, 85, 90). In practice, we observe no evidence of overfitting on the validation
set (see Section 3), but we regard scaling up the held-out set as important future work.</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Discussion</title>
        <p>HTS rulings demand fine-grained reasoning (e.g., partially vs. fully fabricated wafers) and hierarchical
accuracy at 6- and 10-digit levels. Errors have direct compliance costs, making this dataset a realistic
and impactful benchmark for evaluating structured reasoning in large language models.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Model Training</title>
      <p>While several open-source large language models could, in principle, be adapted for tarif classification,
we made a deliberate and principled choice to focus exclusively on LLaMA-3.3-70B [13]. Two factors
motivated this decision. First, practical budget constraints made it infeasible to fine-tune multiple frontier
models at scale. Second, LLaMA-3.3-70B is a dense architecture, making it both simpler to fine-tune
and easier to deploy in inference settings compared to Mixture-of-Experts (MoE) architectures such as
DeepSeek-R1 or GPT-OSS-120B. From a community perspective, providing a dense and reproducible
baseline lowers the entry barrier for downstream research: training and inference pipelines are easier
to set up, memory usage is more predictable, and accuracy is less sensitive to expert routing heuristics.</p>
      <sec id="sec-3-1">
        <title>3.1. Supervised Fine-Tuning Objective</title>
        <p>We adapted LLaMA-3.3-70B to the CROSS dataset using supervised fine-tuning (SFT) [ 14, 15]. Each ruling
was transformed into an input–output pair, where the input is a ruling-derived product description and
the output is the correct HTS code along with a reasoning trace. This makes the task well aligned with
the SFT paradigm, which minimizes the token-level negative log-likelihood of ground-truth outputs.</p>
        <p>Formally, for an input sequence  = (1, . . . , ) and target sequence  = (1, . . . , ), the model
with parameters  defines conditional probabilities   ( | , &lt;). The training loss is then:</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Training Setup and Stability</title>
        <p>Fine-tuning was performed for 5 epochs (approximately 1,400 steps) using the AdamW optimizer with
 1 = 0.9,  2 = 0.95, weight decay = 0.1, and a cosine learning-rate schedule initialized at 1 × 10 −7 . To
manage the high memory footprint of 70B-parameter models, we employed bf16 precision and gradient
accumulation to simulate a batch size of 64 sequences. Training was distributed across 16 × A100-80GB
GPUs using fully sharded data parallelism.</p>
        <p>As shown in Figure 1, the training loss decreases sharply in the first 200 steps and then stabilizes near
convergence, with no sign of overfitting on the validation set. We observed stable gradient norms and
no catastrophic spikes in loss, suggesting that dense models like LLaMA-3.3-70B are well suited to small
but domain-specific datasets when carefully regularized. This highlights that reproducible fine-tuning
of frontier models is feasible even under modest compute budgets, provided that optimization choices
are tuned to stability.
3.3. Beyond Supervised Fine-Tuning: Reinforcement Learning
Although our experiments in this paper are limited to supervised fine-tuning, we view reinforcement
learning (RL) as a promising extension rather than a component of the current Atlas training pipeline.
Concretely, a lightweight and cost-efective starting point would be a rule-based reward model. For
instance, we can define rewards as: 1 when the model correctly predicts the full 10-digit HTS US code,
0.6 when the first 6 digits (globally harmonized HS code) are correct, and −1 otherwise. Formally, for
classification ^ and gold label :</p>
        <p>⎧⎪1,
(^, ) = ⎨</p>
        <p>Such a structured reward can be readily integrated into GRPO [16] or related policy-gradient methods
such as PPO [17]. This approach would allow the model to explore reasoning trajectories that go beyond
memorization, while keeping the reward landscape interpretable and inexpensive to compute.
Importantly, this positions tarif code classification as a promising candidate for lightweight reinforcement
learning research on high-stakes, domain-specific reasoning tasks. We leave the implementation and
empirical evaluation of such RL schemes for HTS classification to future work.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.4. Ablations and Future Work</title>
        <p>While our study focused exclusively on LLaMA-3.3-70B, several ablation studies could provide deeper
insights and further guide the community:
• Model scale: Evaluating smaller LLaMA variants (e.g., 8B or 3B) would clarify the tradeof
between accuracy, cost, and deployability on edge devices.
• Retrieval augmentation: Integrating retrieval over the 17,000-page HTS documents may reduce
hallucinations and improve long-tail classification accuracy, complementing SFT.
• Contrastive and hybrid objectives: Beyond NLL, contrastive learning between closely related
codes (e.g., semiconductor wafers vs. finished chips) may sharpen decision boundaries.
• Direct Preference optimization: Beyond NLL training, methods such as Direct Preference
Optimization (DPO) [18] could leverage structured preferences over HTS classifications (e.g.,
preferring correct 10-digit codes over near-misses, or valid reasoning traces over hallucinated
ones). This would allow the model to learn not just to imitate CROSS rulings but to actively steer
away from incorrect classifications.
• RL scaling studies: Comparing rule-based GRPO with preference-based RLHF could quantify
the cost–benefit tradeofs of reinforcement learning at 70B scale.</p>
        <p>These directions highlight that while Atlas establishes a strong dense-model baseline, HTS
classification remains an open problem with substantial room for methodological innovation.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Evaluation</title>
      <p>
        We evaluate all models on a held-out test set of 200 CROSS rulings, predicting the correct 10-digit HTS
US code per product. Because the classification is hierarchical, we report three metrics: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) full 10-digit
match, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) partial 6-digit match (globally harmonized), and (3) average digits correct (
        <xref ref-type="bibr" rid="ref1 ref2">0–10</xref>
        ).
Average digits correct. Beyond exact 6- and 10-digit accuracy, we report average digits correct,
which measures how many leading digits of the 10-digit HTS US code the model gets right on average.
For a predicted code ^ and gold code , we compute the longest matching prefix length (^, ) ∈
{0, 1, . . . , 10} and then average  over the test set. This metric captures how close near-miss predictions
are along the hierarchical tree: a model that consistently predicts the correct chapter and heading but
misses fine-grained subheadings will have a high average prefix length even if its 10-digit accuracy is
modest.
      </p>
      <sec id="sec-4-1">
        <title>4.1. Accuracy at 10 and 6 Digits</title>
        <p>Table 3 shows fully correct classifications. GPT-5-Thinking 1 achieves 25%, while Atlas attains 40%,
the highest among all models.. At the 6-digit level (Table 3), Atlas also leads with 57.5%, slightly
1Model outputs and pricing were obtained from public API documentation and experiments conducted in September 2025.
above GPT-5’s 55.5%. These results confirm that domain-specific fine-tuning improves both global and
U.S.-specific accuracy.</p>
        <p>Model
GPT-5-Thinking
Gemini-2.5-Pro-Thinking
DeepSeek-R1 (05/28)
GPT-OSS-120B
LLaMA-3.3-70B
Atlas (fine-tuned)
2.5
DeepSeek
1.5</p>
        <p>8
GPT-OSS</p>
        <p>20.7
2.1
LLaMA</p>
        <p>Interestingly, GPT-5-Thinking and Atlas are much closer at the 6-digit (globally harmonized) level
than at the full 10-digit US level (55.5% vs. 57.5%). We view this as evidence that large proprietary models
already capture a substantial amount of generic semantic knowledge about product categories and HS
chapters, while the remaining performance gap at 10 digits is driven by U.S.-specific subheadings and
edge cases that benefit disproportionately from domain-adapted fine-tuning on CROSS rulings. In other
words, Atlas provides most of its value in the last few digits of the hierarchy and in its open-weight,
cost-eficient deployability, rather than simply re-learning globally harmonized structure that frontier
models already approximate.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Inference Cost</title>
        <p>Cost per classification is crucial for scalability. Table 4 shows that open-source models, particularly
Atlas, are an order of magnitude cheaper while maintaining state-of-the-art accuracy.</p>
        <p>Cost estimates for proprietary models are based on public API pricing as of September 2025 for
prompt and completion tokens at the context lengths used in our experiments (see footnote in Section 4).
For open-weight models, including the base LLaMA-3.3-70B and Atlas, we report an approximate
per-1,000 inference cost assuming on-premise deployment on A100-80GB GPUs with standard cloud
pricing. While absolute numbers will vary with hardware, batch size, and provider, the relative ordering
(frontier APIs being several times more expensive than open-weight deployment at similar throughput)
is robust.</p>
        <p>Model
GPT-5-Thinking
Gemini-2.5-Pro-Thinking
DeepSeek-R1
GPT-OSS-120B
LLaMA-3.3-70B
Atlas (fine-tuned)</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Discussion</title>
        <p>Taken together, these results highlight a critical tradeof: Atlas not only surpasses GPT-5-Thinking
in accuracy (40% vs. 25% fully correct classifications), but also reduces inference cost by nearly 5 ×
compared to GPT-5 and almost 8× compared to Gemini-2.5-Pro-Thinking. Moreover, the strong
performance on partially correct classifications demonstrates that Atlas generalizes beyond U.S.-specific
tarifs to the globally harmonized 6-digit regime, reinforcing its utility for international trade applications.
A qualitative comparison illustrating model reasoning diferences is provided in Appendix C.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Summary and Future Directions</title>
      <p>We introduced the first benchmark for Harmonized Tarif Schedule (HTS) code classification and
presented Atlas, a fine-tuned LLaMA-3.3-70B model for global trade compliance. The study establishes
HTS classification as a challenging new LLM benchmark, with three main takeaways:
• Performance: Atlas achieves 40% fully correct and 57.5% partially correct (6-digit) classifications,
surpassing GPT-5-Thinking (+15 pts) and Gemini-2.5-Pro (+27.5 pts).
• Eficiency: Atlas is 5 × cheaper than GPT-5 and 8× cheaper than Gemini, while supporting
secure self-hosted deployment.
• Challenge: Even the best model attains only 40% 10-digit accuracy, underscoring substantial
headroom for progress.</p>
      <p>Future work includes expanding coverage beyond semiconductors, distilling Atlas into smaller
(8B–3B) variants for edge use, and applying reinforcement learning via rule-based rewards [16, 18] to
improve reasoning and alignment.</p>
      <p>We release Atlas as an open-weight model, allowing organizations to self-host the model under
their own compliance and privacy constraints.</p>
    </sec>
    <sec id="sec-6">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the author(s) used large language models (e.g., OpenAI GPT-4 and
GPT-5) for grammar checking, wording refinement, technical editing, and assistance in restructuring
some paragraphs. No generative AI tools were used to create figures, tables, or experimental results.
After using these tools, the author(s) reviewed, verified, and edited all generated content as needed and
take full responsibility for the publication’s final text and accuracy.
optimization: Your language model is secretly a reward model, Advances in Neural Information
Processing Systems 36 (2023).</p>
    </sec>
    <sec id="sec-7">
      <title>A. Full Dataset Distribution</title>
    </sec>
    <sec id="sec-8">
      <title>B. Prompt Template for Data Transformation</title>
      <p>Each ruling was converted into a structured instruction–response pair to facilitate supervised fine-tuning
of language models. The complete transformation prompt is shown below.</p>
      <p>Given the following HTS ruling information:
HTS Code: {hts_code}
Ruling Number: {ruling_number}
Title: {title}
Date: {date}
URL: {url}
Summary: {summary}
Content: {content}
Please analyze this information and provide:
a) A concise product description representing the item being classified
b) A reasoning path justifying why the HTS US code is correct
c) The final HTS US code
Format your response as follows:
User: What is the HTS US Code for [product_description]?
Model:
HTS US Code -&gt; [HTS US Code]
Reasoning -&gt; [detailed_reasoning_path]</p>
      <p>This design enforces explicit reasoning traces, aligning with recent advances in chain-of-thought
modeling [12].</p>
    </sec>
    <sec id="sec-9">
      <title>C. Qualitative Example</title>
      <p>Ground
HTS Code
Atlas (Ours)
GPT-5-Thinking</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>United</given-names>
            <surname>States International Trade Commission</surname>
          </string-name>
          ,
          <article-title>Harmonized tarif schedule (hts us)</article-title>
          , https://hts. usitc.gov/,
          <year>2025</year>
          . Accessed:
          <fpage>2025</fpage>
          -09-20.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>Gemini-2</source>
          .
          <fpage>5</fpage>
          -Pro- Reasoning:
          <article-title>Associates with silicon materials but ignores doping context</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Thinking</given-names>
            <surname>Prediction</surname>
          </string-name>
          :
          <volume>3824</volume>
          .
          <fpage>99</fpage>
          .99.99 (
          <issue>×</issue>
          )
          <article-title>Incorrect; generic chemical compound</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>