Triangle104
/

Llama-3.1-Nemotron-Nano-8B-v1-Q8_0-GGUF

Text Generation

Model card Files Files and versions Community

Triangle104 commited on 7 days ago

Commit

8441d43

·

verified ·

1 Parent(s): c9c970f

Update README.md

Files changed (1) hide show

README.md +24 -0

README.md CHANGED Viewed

@@ -19,6 +19,30 @@ tags:
 This model was converted to GGUF format from [`nvidia/Llama-3.1-Nemotron-Nano-8B-v1`](https://huggingface.co/nvidia/Llama-3.1-Nemotron-Nano-8B-v1) using llama.cpp via the ggml.ai's [GGUF-my-repo](https://huggingface.co/spaces/ggml-org/gguf-my-repo) space.
 Refer to the [original model card](https://huggingface.co/nvidia/Llama-3.1-Nemotron-Nano-8B-v1) for more details on the model.
 ## Use with llama.cpp
 Install llama.cpp through brew (works on Mac and Linux)

 This model was converted to GGUF format from [`nvidia/Llama-3.1-Nemotron-Nano-8B-v1`](https://huggingface.co/nvidia/Llama-3.1-Nemotron-Nano-8B-v1) using llama.cpp via the ggml.ai's [GGUF-my-repo](https://huggingface.co/spaces/ggml-org/gguf-my-repo) space.
 Refer to the [original model card](https://huggingface.co/nvidia/Llama-3.1-Nemotron-Nano-8B-v1) for more details on the model.
+---
+Llama-3.1-Nemotron-Nano-8B-v1 is a large language model (LLM) which is a derivative of Meta Llama-3.1-8B-Instruct
+ (AKA the reference model). It is a reasoning model that is post trained
+ for reasoning, human chat preferences, and tasks, such as RAG and tool
+calling.
+Llama-3.1-Nemotron-Nano-8B-v1 is a model which offers a great
+tradeoff between model accuracy and efficiency. It is created from Llama
+ 3.1 8B Instruct and offers improvements in model accuracy. The model
+fits on a single RTX GPU and can be used locally. The model supports a
+context length of 128K.
+This model underwent a multi-phase post-training process to enhance
+both its reasoning and non-reasoning capabilities. This includes a
+supervised fine-tuning stage for Math, Code, Reasoning, and Tool Calling
+ as well as multiple reinforcement learning (RL) stages using REINFORCE
+(RLOO) and Online Reward-aware Preference Optimization (RPO) algorithms
+for both chat and instruction-following. The final model checkpoint is
+obtained after merging the final SFT and Online RPO checkpoints.
+Improved using Qwen.
+---
 ## Use with llama.cpp
 Install llama.cpp through brew (works on Mac and Linux)