Instructions to use VishwamAI/vishwamai-merged-gemma with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use VishwamAI/vishwamai-merged-gemma with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="VishwamAI/vishwamai-merged-gemma") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("VishwamAI/vishwamai-merged-gemma") model = AutoModelForMultimodalLM.from_pretrained("VishwamAI/vishwamai-merged-gemma", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use VishwamAI/vishwamai-merged-gemma with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "VishwamAI/vishwamai-merged-gemma" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VishwamAI/vishwamai-merged-gemma", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/VishwamAI/vishwamai-merged-gemma
- SGLang
How to use VishwamAI/vishwamai-merged-gemma with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "VishwamAI/vishwamai-merged-gemma" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VishwamAI/vishwamai-merged-gemma", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "VishwamAI/vishwamai-merged-gemma" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VishwamAI/vishwamai-merged-gemma", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Desktop
- Docker Model Runner
How to use VishwamAI/vishwamai-merged-gemma with Docker Model Runner:
docker model run hf.co/VishwamAI/vishwamai-merged-gemma
VishwamAI — Merged fp16 Model
A safe, local cybersecurity + math/science reasoning assistant, fine-tuned from Gemma 4 E4B-it (Apache 2.0). This repository holds the merged fp16 weights — a complete, standalone model (no adapter or base download required).
⚠️ Scope & authorized use only. VishwamAI is a defensive-first assistant. Use its security capabilities only against systems you own or are explicitly authorized to test (labs, CTFs, defensive research, education). It is not intended for attacking third parties or any unlawful activity.
What it is
VishwamAI is a 4B-class assistant designed to run locally on modest hardware while handling four kinds of work:
- 🛡️ Defensive security — MITRE ATT&CK explanation, detection engineering (Sysmon/Sigma/YARA), DFIR & log/pcap triage, secure-code review, CTF methodology.
- 🔢 Math & science — step-by-step reasoning where exact computation is delegated to a sandboxed SymPy/units tool (the model reasons, a CAS computes → verifiable answers).
- 💻 Coding — writes code that is verified by running it in a locked-down sandbox.
- 💬 General chat & reasoning — a general slice was kept in training to limit catastrophic forgetting.
Design principle: reason with the model, verify with a tool. Facts (CVEs, ATT&CK IDs, constants) should be served from a local RAG index rather than trusted from the weights.
Model details
| Base model | unsloth/gemma-4-E4B-it (Apache 2.0) |
| Method | QLoRA SFT (Unsloth + TRL), rank 16, merged to fp16 |
| Context | up to 128K (base); trained at 2048 |
| Precision | fp16 safetensors (sharded) |
| Language | English |
How to run
A) Transformers (GPU with enough VRAM)
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "VishwamAI/vishwamai-merged"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.float16, device_map="auto")
msgs = [{"role": "user", "content": "Explain MITRE ATT&CK T1003.001 and give two Sysmon detections."}]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(input_ids=inputs, max_new_tokens=300, temperature=0.7)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
B) Local GGUF (recommended for laptops — llama.cpp / Ollama)
Convert the fp16 weights to a small quantized GGUF on a machine with ≥16 GB RAM:
huggingface-cli download VishwamAI/vishwamai-merged --local-dir ./vishwamai-merged
python convert_hf_to_gguf.py ./vishwamai-merged \
--outfile vishwamai-f16.gguf --outtype f16 --use-temp-file
./llama-quantize vishwamai-f16.gguf vishwamai-Q4_K_M.gguf Q4_K_M
# serve an OpenAI-compatible API on localhost
./llama-server -m vishwamai-Q4_K_M.gguf -ngl 99 -c 8192 --host 127.0.0.1 --port 8080
Then call it:
import requests
r = requests.post("http://127.0.0.1:8080/v1/chat/completions", json={
"model": "vishwamai", "temperature": 0.7,
"messages": [{"role": "user", "content": "Give two Sysmon detections for LSASS dumping."}]})
print(r.json()["choices"][0]["message"]["content"])
Recommended system prompt
You are VishwamAI, a defensive-first security research assistant and a
math/science solver for an authorized practitioner. Be precise, state uncertainty,
never fabricate CVE IDs or physical constants, and use the calculator/CAS tool for
any non-trivial computation.
Training data (summary)
A mixed instruction set (~security ~40–50%, math/science ~20–25%, general ~15–20%, plus ~3–5% safety refusals and ~5–10% contrastive "allowed defensive" examples to reduce over-refusal). Public sources used may include Trend Micro Primus-Instruct (ODC-By), MITRE ATT&CK, and NVD/CVE-derived reasoning tasks, alongside the author's own notes. Benchmark items were held out of training.
Safety
Fine-tuning can erode alignment, so safety is treated as a system: curated refusals + contrastive allowed examples in training, an optional runtime guard model (e.g. Llama Guard / Qwen3Guard) on outputs, prompt-injection handling that treats all retrieved/tool content as untrusted data, and sandboxed tool execution (no host shell). Re-run a safety eval after every retrain/re-quantization.
Limitations
- ~4B model — capable but weaker than frontier models on deep open-ended reasoning.
- May hallucinate facts (CVE IDs, constants) — route factual lookups through RAG and keep the "state uncertainty" system prompt.
- Quantization shifts behavior — evaluate the GGUF you actually deploy.
- Guard models are imperfect and tuned for general consumer harms, not security nuance.
License
Released under Apache 2.0, inheriting the license of the Gemma 4 base. Review the licenses/terms of any datasets you add before redistribution.
Disclaimer
Provided for authorized security research, education, and lawful use only. The authors accept no liability for misuse. You are responsible for complying with all applicable laws and for having authorization to test any system.
- Downloads last month
- 34