All You Need To Know About Running LLMs Locally @bycloudAI
All You Need To Know About Running LLMs Locally  @bycloudAI
Uploaded February 2024 | Updated September 2026, 1 week ago
my latest project: Intuitive AI Academy, learn modern AI/LLMs Intuitively
https://intuitiveai.academy/
code "NYNM" for 50% off forever (limited to 50)


TensorRT LLM
[Code] github.com/NVIDIA/TensorRT-LLM
[Getting Started Blog] nvda.ws/3O7f8up
[Dev Blog] nvda.ws/490uadi

Chat with RTX
[Download] nvda.ws/3OHPRHE
[Blog] nvda.ws/3whKZTb

Links:
[Oobabooga] github.com/oobabooga/text-generation-webui
[SillyTavern] github.com/SillyTavern/SillyTavern
[LM Studio] lmstudio.ai
[Axolotl] github.com/OpenAccess-AI-Collective/axolotl
[Llama Factory] github.com/hiyouga/LLaMA-Factory
[HuggingFace] huggingface.co/models
[AWQ] github.com/mit-han-lab/llm-awq
[ExLlamav2] github.com/turboderp/exllamav2
[GGUF] github.com/ggerganov/ggml/blob/master/docs/gguf.md
[GPTQ] github.com/IST-DASLab/gptq
[LlamaCpp] github.com/ggerganov/llama.cpp
[vllm] github.com/vllm-project/vllm
[TensorRT LLM] github.com/NVIDIA/TensorRT-LLM
[Chat with RTX] nvidia.com/en-us/ai-on-rtx/chat-with-rtx-generative-ai
[LlamaIndex] github.com/run-llama/llama_index
[Continue.dev] https://continue.dev/

Model recommendations (I know you are here after DeepSeek):
[All DeepSeek Models] huggingface.co/collections/deepseek-ai/deepseek-r1-678e1e131c0169c0bc89728d
[Easily Download with Ollama] ollama.com/library/deepseek-r1
Here's the rule of thumb to know if you can run it:
If your VRAM is larger than the model GB size * 1.2, than you can run that model size locally.
Eg. DeepSeek-7B = 4.7GB then 4.7*1.2=5.64, so if your GPU has 8GB VRAM, since 8GB is bigger than 5.64, you can run DeepSeek-7B.
Check out my latest video on DeepSeek-R1 to understand the context better!

(the following are all outdated)
Just use Llama-3.1 instead for everything.
[Llama-3.1] huggingface.co/collections/meta-llama/llama-31-669fc079a0c406a149a5738f
Translation can try Aya 23
[Aya 23] huggingface.co/CohereForAI/aya-23-8B

(the following are all outdated)
[Nous-Hermes-llama-2-7b] huggingface.co/NousResearch/Nous-Hermes-llama-2-7b
[Openchat-3.5-0106] huggingface.co/openchat/openchat-3.5-0106
[SOLAR-10.7B-Instruct-v1.0] huggingface.co/upstage/SOLAR-10.7B-Instruct-v1.0
[Google Gemma] huggingface.co/google/gemma-7b
[Mixtral-8x7B-Instruct-v0.1] huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1
[Deepseek-coder-33b-instruct] huggingface.co/deepseek-ai/deepseek-coder-33b-instruct
[Colbertv2.0] huggingface.co/colbert-ir/colbertv2.0


This video is supported by the kind Patrons & YouTube Members:
🙏Andrew Lescelius, alex j, Chris LeDoux, Alex Maurice, Miguilim, Deagan, FiFaŁ, Daddy Wen, Tony Jimenez, Panther Modern, Jake Disco, Demilson Quintao, Shuhong Chen, Hongbo Men, happi nyuu nyaa, Carol Lo, Mose Sakashita, Miguel, Bandera, Gennaro Schiano, gunwoo, Ravid Freedman, Mert Seftali, Mrityunjay, Richárd Nagyfi, Timo Steiner, Henrik G Sundt, projectAnthony, Brigham Hall, Kyle Hudson, Kalila, Jef Come, Jvari Williams, Tien Tien, BIll Mangrum, owned, Janne Kytölä, SO, Richárd Nagyfi, Hector, Drexon

[Discord] discord.gg/NhJZGtH
[Twitter] twitter.com/bycloudai
[Patreon] patreon.com/bycloud

[Music] massobeats - magic carousel
[Profile & Banner Art] twitter.com/pygm7
[Video Editor] maikadihaika
All You Need To Know About Running LLMs LocallyThe Difference of AI Videos No One Tells You AboutAttention Sink: The Fluke That Made LLMs Actually UsableThe AI Hardware Arms Race Is Getting Out of HandGoogles New OS Gemma 4 Series Beats Models 10x Its Size!Has ChatGPT Finally Been Dethroned? Claude 3 ReviewThis is why you should be nice to LLMs... #AI #Claude #LLM1,000,000,000 Tokens LLM, SDXL Leaked, Midjourney WEIRD & More [The AI Timeline #8]How did a 27M Model even beat ChatGPT?AI News: DeepSeek-R1 V2, the new open source SoTA!Gemini 1.5 Pro With 10,000,000 Tokens Is AbsurdMicrosoft Just Dropped LLMs Frontier Data Engineering Secrets
bycloud |

All You Need To Know About Running LLMs Locally

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER