stabilityai
/

stablelm-zephyr-3b

Text Generation

Transformers

Safetensors

Model card Files Files and versions Community

jon-tow

leaderboard-pr-bot commited on Mar 7

Commit

015f44c

•

1 Parent(s): 8b471c7

Adding Evaluation Results (#14)

Browse files

- Adding Evaluation Results (ba1eb10abfbc233fa99d72aeadb9f6dd81359ec2)

Co-authored-by: Open LLM Leaderboard PR Bot <[email protected]>

Files changed (1) hide show

README.md +122 -6

README.md CHANGED Viewed

@@ -1,21 +1,124 @@
 ---
 datasets:
 - HuggingFaceH4/ultrachat_200k
 - HuggingFaceH4/ultrafeedback_binarized
 - meta-math/MetaMathQA
 - WizardLM/WizardLM_evol_instruct_V2_196k
 - Intel/orca_dpo_pairs
-language:
-- en
-tags:
-- causal-lm
 extra_gated_fields:
   Name: text
   Email: text
   Country: text
   Organization or Affiliation: text
   I ALLOW Stability AI to email me about new model releases: checkbox
-license: other
 ---
 # `StableLM Zephyr 3B`
@@ -150,4 +253,17 @@ The model is intended to be used as a foundational base model for application-sp
 This model is not trained against adversarial inputs. We strongly recommend pairing this model with an input and output classifier to prevent harmful responses.
-Through our internal red teaming, we discovered that while the model will not output harmful information if not prompted to do so, it is willing to output potentially harmful outputs or misinformation when the user requests it. Using this model will require guardrails around your inputs and outputs to ensure that any outputs returned are not misinformation or harmful. Additionally, as each use case is unique, we recommend running your own suite of tests to ensure proper performance of this model. Finally, do not use the models if they are unsuitable for your application, or for any applications that may cause deliberate or unintentional harm to others.

 ---
+language:
+- en
+license: other
+tags:
+- causal-lm
 datasets:
 - HuggingFaceH4/ultrachat_200k
 - HuggingFaceH4/ultrafeedback_binarized
 - meta-math/MetaMathQA
 - WizardLM/WizardLM_evol_instruct_V2_196k
 - Intel/orca_dpo_pairs
 extra_gated_fields:
   Name: text
   Email: text
   Country: text
   Organization or Affiliation: text
   I ALLOW Stability AI to email me about new model releases: checkbox
+model-index:
+- name: stablelm-zephyr-3b
+  results:
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: AI2 Reasoning Challenge (25-Shot)
+      type: ai2_arc
+      config: ARC-Challenge
+      split: test
+      args:
+        num_few_shot: 25
+    metrics:
+    - type: acc_norm
+      value: 46.08
+      name: normalized accuracy
+    source:
+      url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=stabilityai/stablelm-zephyr-3b
+      name: Open LLM Leaderboard
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: HellaSwag (10-Shot)
+      type: hellaswag
+      split: validation
+      args:
+        num_few_shot: 10
+    metrics:
+    - type: acc_norm
+      value: 74.16
+      name: normalized accuracy
+    source:
+      url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=stabilityai/stablelm-zephyr-3b
+      name: Open LLM Leaderboard
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: MMLU (5-Shot)
+      type: cais/mmlu
+      config: all
+      split: test
+      args:
+        num_few_shot: 5
+    metrics:
+    - type: acc
+      value: 46.17
+      name: accuracy
+    source:
+      url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=stabilityai/stablelm-zephyr-3b
+      name: Open LLM Leaderboard
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: TruthfulQA (0-shot)
+      type: truthful_qa
+      config: multiple_choice
+      split: validation
+      args:
+        num_few_shot: 0
+    metrics:
+    - type: mc2
+      value: 46.49
+    source:
+      url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=stabilityai/stablelm-zephyr-3b
+      name: Open LLM Leaderboard
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: Winogrande (5-shot)
+      type: winogrande
+      config: winogrande_xl
+      split: validation
+      args:
+        num_few_shot: 5
+    metrics:
+    - type: acc
+      value: 65.51
+      name: accuracy
+    source:
+      url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=stabilityai/stablelm-zephyr-3b
+      name: Open LLM Leaderboard
+  - task:
+      type: text-generation
+      name: Text Generation
+    dataset:
+      name: GSM8k (5-shot)
+      type: gsm8k
+      config: main
+      split: test
+      args:
+        num_few_shot: 5
+    metrics:
+    - type: acc
+      value: 42.15
+      name: accuracy
+    source:
+      url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=stabilityai/stablelm-zephyr-3b
+      name: Open LLM Leaderboard
 ---
 # `StableLM Zephyr 3B`
 This model is not trained against adversarial inputs. We strongly recommend pairing this model with an input and output classifier to prevent harmful responses.
+Through our internal red teaming, we discovered that while the model will not output harmful information if not prompted to do so, it is willing to output potentially harmful outputs or misinformation when the user requests it. Using this model will require guardrails around your inputs and outputs to ensure that any outputs returned are not misinformation or harmful. Additionally, as each use case is unique, we recommend running your own suite of tests to ensure proper performance of this model. Finally, do not use the models if they are unsuitable for your application, or for any applications that may cause deliberate or unintentional harm to others.
+# [Open LLM Leaderboard Evaluation Results](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard)
+Detailed results can be found [here](https://huggingface.co/datasets/open-llm-leaderboard/details_stabilityai__stablelm-zephyr-3b)
+|             Metric              |Value|
+|---------------------------------|----:|
+|Avg.                             |53.43|
+|AI2 Reasoning Challenge (25-Shot)|46.08|
+|HellaSwag (10-Shot)              |74.16|
+|MMLU (5-Shot)                    |46.17|
+|TruthfulQA (0-shot)              |46.49|
+|Winogrande (5-shot)              |65.51|
+|GSM8k (5-shot)                   |42.15|