All posts

Research Note: Large Language Models and Thinking Tokens

Dr. Jerry A. Smith · February 14, 2025 · 7 min read

Large Language Models (LLMs) have revolutionized natural language processing, but their ability to reason and solve complex problems remains an active area of research. This report explores “thinking tokens,” a novel approach to enhance LLMs’ reasoning capabilities. By inserting unique tokens into the model’s input, researchers hope to give LLMs more time to process information and perform calculations internally before generating responses. We examine the implementation, benefits, and challenges of thinking tokens and their potential impact on the future of AI reasoning and problem-solving.

Tokens: The Building Blocks of LLM Language Processing

Tokenization is a critical yet often overlooked process that fundamentally shapes how large language models understand and generate text. Tokens are the basic units LLMs process, typically representing words, parts of words, or individual characters. The choice of tokenization algorithm can significantly impact model performance and efficiency.

Modern LLMs use subword tokenization methods like Byte Pair Encoding (BPE). This approach finds common pairs of characters in the training data and compresses them into single tokens, balancing vocabulary size with efficient text representation. For example, the GPT-2 tokenizer might split “tokenization” into separate tokens like “token,” “iza,” and “tion.”

Tokenization presents several challenges:

  • Inconsistent handling across languages
  • Inefficient processing of special characters and whitespace
  • Arbitrary splitting of words based on capitalization or position

These issues can lead to poor performance on arithmetic, spelling, and handling non-English text. However, recent research challenges the fundamental assumptions of tokenization. The T-FREE approach maps words directly to sparse patterns based on character sequences, potentially reducing model size by 85% while maintaining performance. This demonstrates that questioning core assumptions can lead to significant LLM design and efficiency breakthroughs.

Sources

Introduction to Thinking Tokens in LLMs

Thinking tokens represent a novel approach to enhancing reasoning capabilities in large language models by allowing for internal computation before generating responses. Unlike traditional chain-of-thought methods that rely on verbalized reasoning steps, thinking tokens operate in the model’s latent space, enabling more efficient and flexible processing of complex logical problems.

Herel and Mikolov introduced the concept in their 2023 paper, which proposed inserting special “thinking tokens” after each word in a sentence when encountering complex problems. This approach aims to give the model more time to perform calculations internally before producing an output.

Recent research by Meta FAIR and UC San Diego has expanded on this idea with their COCONUT model (Chain Of CONtinUous Thought). COCONUT encodes “latent thoughts” that replace individual reasoning steps, allowing the model to explore multiple potential paths simultaneously without converting to and from natural language at each step.

While thinking tokens have shown promise for improving performance on specific logical reasoning tasks, their effectiveness for general instruction-following remains an active research and development area.

Sources

Reinforcement Learning in Large Language Models

**Reinforcement learning (RL) is crucial for improving reasoning in large language models (LLMs), complementing supervised fine-tuning (SFT) to enhance accuracy, consistency, and response clarity.** Techniques like Odds Ratio Preference Optimization (ORPO) and Group Relative Policy Optimization (GRPO) address specific LLM reasoning and consistency challenges. For example, ORPO combines cross-entropy loss with preference optimization, adjusting model probabilities to favor preferred answers while penalizing rejected ones. This approach improves consistency, as evidenced by higher Majority@K scores in experiments.

However, challenges persist in enhancing Pass@K scores, which measure the generation of novel correct answers. Researchers are exploring advanced RL methods like GRPO and techniques that encourage self-correction, such as “wait” prompts. Scaling experiments to larger models and more complex datasets offers another promising avenue for overcoming limitations.

Tools like SG Lang and parameter-efficient fine-tuning methods such as Low-Rank Adaptation (LoRA) support efficient training. These approaches enable optimization without extensive computational resources, which is crucial for advancing LLM capabilities in reasoning tasks.

Sources

Analysis of Thinking Tokens in LLMs

Thinking tokens provide marginal improvements in LLM reasoning capabilities but consistently underperform compared to Chain-of-Thought (CoT) approaches. While conceptually appealing, studies show thinking tokens offer only slight performance gains across multiple benchmarks. This underperformance likely stems from a reliance on a single reasoning path, in contrast to CoT methods that explore numerous potential solutions.

A key advantage of thinking tokens is their ability to facilitate step-by-step reasoning without requiring model fine-tuning. This allows for more transparent intermediate steps in the problem-solving process. However, the token overhead introduced by thinking tokens can significantly increase computational costs without proportional accuracy improvements.

Recent research has explored hybrid approaches combining thinking tokens with other techniques. For example, integrating thinking tokens into retrieval-augmented models shows promise for enhancing reasoning on knowledge-intensive tasks. Additionally, some studies have experimented with dynamic token allocation based on problem complexity to optimize the trade-off between performance and efficiency.

Despite limitations, thinking tokens remain an active area of research as the field seeks more robust reasoning methods for LLMs. Future work may focus on developing more sophisticated token allocation strategies or combining thinking tokens with complementary reasoning approaches.

Sources

Drawbacks and Limitations of Thinking Tokens in LLMs

Thinking tokens may not constantly improve model performance and can introduce new challenges. While the concept aims to give language models more time to process complex problems, experiments show mixed results. Adding thinking tokens after each word in a dataset led to slight decreases in perplexity across standard language modeling tasks. This suggests that more “thinking time” does not necessarily translate to better outputs.

One key limitation is the potential for increased computational costs. Each additional thinking token expands the input sequence, requiring more processing power and memory. This could negate efficiency gains, especially for resource-constrained applications.

Furthermore, thinking tokens may disrupt the natural flow of language. A case study on mathematical reasoning found that while thinking tokens improved performance on some complex calculations, they also increased the likelihood of the model losing context or generating irrelevant outputs when overused.

The optimal number and placement of thinking tokens remain an open research question. Too few may not provide sufficient processing time, while too many risks overwhelm the model’s ability to maintain coherence across longer sequences.

Sources

Practical Applications of LLMs with Thinking Tokens

Thinking Tokens show promise for enhancing LLM reasoning but currently underperform compared to Chain-of-Thought prompting. TTs aim to facilitate unsupervised reasoning by inserting intermediate “thinking” tokens, allowing models more time to compute latent states before generating output. However, empirical results demonstrate that TTs only marginally improve performance and consistently fall short of CoT across multiple benchmarks.

A key limitation is TTs’ reliance on a single embedding, which can introduce noisy gradients and inconsistent learning signals. This is particularly problematic for tasks requiring structured intermediate steps, like arithmetic reasoning or multi-hop commonsense tasks.

Researchers are exploring ways to address these challenges. For example, using multiple distinct thinking token embeddings shows promise in producing clearer gradients and more expressive representations. Additionally, combining TTs with CoT approaches may leverage the strengths of both techniques.

While current practical applications remain limited, TTs represent an important area of ongoing research for improving unsupervised reasoning capabilities in LLMs.

Sources

Speculation on the Future of Thinking Tokens in LLMs

Thinking Tokens show promise for enhancing LLM reasoning capabilities, but face challenges in implementation and effectiveness. While designed to extend reasoning by inserting “thinking” time, Thinking Tokens (TTs) have underperformed compared to Chain-of-Thought (CoT) prompting across multiple benchmarks. This may stem from reliance on a single embedding, resulting in inconsistent learning signals and noisy gradients.

Research suggests improvements by using multiple distinct TT embeddings rather than a single shared token. For example, experiments with two unique TT embeddings resulted in clearer gradients and more dynamic embedding movement during training compared to a single TT.

Future directions for TT development may include:

  • Exploring variable-length TT sequences
  • Integrating TTs with other techniques like retrieval-augmented generation
  • Developing TT architectures optimized for specific reasoning tasks

Ultimately, while TTs offer an intriguing unsupervised approach to reasoning, significant refinement is still needed before they can match or exceed the performance of structured prompting techniques like CoT. Ongoing research into TT optimization and integration with other LLM advancements will be crucial for realizing their full potential.

Sources

Summary and Future Outlook

Thinking tokens are innovative in enhancing reasoning capabilities in large language models, but their practical impact remains limited. While conceptually promising, studies consistently show that thinking tokens underperform over established Chain-of-Thought (CoT) methods across various benchmarks. The primary challenges include:

* Reliance on single embeddings leading to noisy gradients
* Potential disruption of natural language flow
* Increased computational costs without proportional accuracy gains

Despite these limitations, ongoing research explores promising avenues for improvement:

* Utilizing multiple distinct thinking token embeddings
* Combining thinking tokens with retrieval-augmented models
* Developing dynamic token allocation strategies

As AI continues to evolve, the concept of thinking tokens may play a role in future advancements. However, significant refinement is needed before they can match or surpass the effectiveness of current reasoning techniques in large language models.

Start with one workflow.

Tell me what your team does today, where the work gets stuck, and what a useful result would look like. We will use a short call to identify a sensible next step.

Book a call