dandelin
/

vilt-b32-mlm

Inference Endpoints

Model card Files Files and versions Community

nielsr HF staff commited on Jan 23, 2022

Commit

7051201

•

1 Parent(s): 45b9dfe

Add code example

Files changed (1) hide show

README.md +46 -6

README.md CHANGED Viewed

@@ -9,17 +9,57 @@ Vision-and-Language Transformer (ViLT) model pre-trained on GCC+SBU+COCO+VG (200
 Disclaimer: The team releasing ViLT did not write a model card for this model so this model card has been written by the Hugging Face team.
-## Model description
-(to do)
 ## Intended uses & limitations
-You can use the raw model for visual question answering.
 ### How to use
-(to do)
 ## Training data

 Disclaimer: The team releasing ViLT did not write a model card for this model so this model card has been written by the Hugging Face team.
 ## Intended uses & limitations
+You can use the raw model for masked language modeling given an image and a piece of text with [MASK] tokens.
 ### How to use
+Here is how to use this model in PyTorch:
+```
+from transformers import ViltProcessor, ViltForMaskedLM
+import requests
+from PIL import Image
+import re
+url = "http://images.cocodataset.org/val2017/000000039769.jpg"
+image = Image.open(requests.get(url, stream=True).raw)
+text = "a bunch of [MASK] laying on a [MASK]."
+processor = ViltProcessor.from_pretrained("dandelin/vilt-b32-mlm")
+model = ViltForMaskedLM.from_pretrained("dandelin/vilt-b32-mlm")
+# prepare inputs
+encoding = processor(image, text, return_tensors="pt")
+# forward pass
+outputs = model(**encoding)
+tl = len(re.findall("\[MASK\]", text))
+inferred_token = [text]
+# gradually fill in the MASK tokens, one by one
+with torch.no_grad():
+    for i in range(tl):
+        encoded = processor.tokenizer(inferred_token)
+        input_ids = torch.tensor(encoded.input_ids).to(device)
+        encoded = encoded["input_ids"][0][1:-1]
+        outputs = model(input_ids=input_ids, pixel_values=pixel_values)
+        mlm_logits = outputs.logits[0]  # shape (seq_len, vocab_size)
+        # only take into account text features (minus CLS and SEP token)
+        mlm_logits = mlm_logits[1 : input_ids.shape[1] - 1, :]
+        mlm_values, mlm_ids = mlm_logits.softmax(dim=-1).max(dim=-1)
+        # only take into account text
+        mlm_values[torch.tensor(encoded) != 103] = 0
+        select = mlm_values.argmax().item()
+        encoded[select] = mlm_ids[select].item()
+        inferred_token = [processor.decode(encoded)]
+selected_token = ""
+encoded = processor.tokenizer(inferred_token)
+processor.decode(encoded.input_ids[0], skip_special_tokens=True)
+```
 ## Training data