Nemotron-4-340B-Instruct is an English-language chat model optimized for synthetic data generation. This large language model (LLM) is a fine-tuned version of Nemotron-4-340B-Base, designed for single and multi-turn chat use-cases with a 4,096 token context length.
The base model was pre-trained on 9 trillion tokens from diverse English texts, 50+ natural languages, and 40+ coding languages. The instruct model underwent additional alignment steps:
The alignment process used approximately 20K human-annotated samples, while 98% of the data for fine-tuning was synthetically generated. Detailed information about the synthetic data generation pipeline is available in the technical report(opens in new tab).
Modalities
Context
4K
Released
Jun 23, 2024
Knowledge Cutoff
Jun 2023
Token volume and request traffic to this model over time.
Nemotron-4-340B-Instruct is an English-language chat model optimized for synthetic data generation. This large language model (LLM) is a fine-tuned version of Nemotron-4-340B-Base, designed for single and multi-turn chat use-cases with a 4,096 token context length. The base model was pre-trained on 9 trillion tokens from diverse English texts, 50+ natural languages, and 40+ coding languages.
Nemotron-4 340B Instruct has a 4,096 token context window.
Nemotron 3.5 Lightning, Nemotron 3.5 Content Safety (free), Nemotron 3 Ultra and 5 more are other text models from Nvidia.
Nemotron-4 340B Instruct was released on June 23, 2024. Its knowledge cutoff is June 30, 2023.