Run AI Models Locally with LM Studio
Learn how to install LM Studio, choose the right model for your hardware, and use offline AI: requirements, quantization, and getting started.

On this page
Setting realistic expectations. Models running on a home computer are less capable than those used by major online services. For summarizing, paraphrasing, classifying, translating, and reasoning over your documents, they work great; however, for complex tasks, you will notice the difference compared to ChatGPT or Claude.
As of July 2025, LM Studio is also free for business use, with no licenses required.
Requirements #
| System | Requirements |
|---|---|
| macOS | Apple Silicon only, macOS 14 or later; 16 GB RAM recommended |
| Windows | x64 with AVX2 instructions or ARM; 16 GB RAM and 4 GB VRAM recommended |
| Linux | Ubuntu 20.04 or later, distributed as an AppImage for x64 and ARM64 |
It works with 8 GB of memory, but only with the smallest models. The real jump happens at 16 GB, and with 32 GB, you can run almost anything that makes sense on a personal computer.
On Mac with Apple silicon memory is shared between the processor and graphics, which is a huge advantage: a Mac with 32 GB can load models that would require a high-end graphics card on a PC. LM Studio on Mac also uses models in MLX format, which are optimized for those chips and are usually faster. On a PC with an NVIDIA card, make sure to keep your drivers up to date.
How much memory do you need for each model #
That is the key question, and the answer depends on quantization: a technique that reduces the precision of the numbers within the model, making the file much lighter at the cost of a slight drop in quality.
A model is identified by its number of parameters (7B, 14B, 32B, where B stands for billions) and its quantization level, often written as Q4_K_M or similar. The lower the number after the Q, the lighter the file and the more the quality decreases.
For practical reference:
| Model | 4-bit quantized file | Recommended memory |
|---|---|---|
| 3B | approx. 2 GB | 8 GB |
| 7-8B | approx. 4-5 GB | 16 GB |
| 14B | approx. 8-9 GB | 16-24 GB |
| 32B | approx. 18-20 GB | 32 GB |
| 70B | approx. 40 GB | 64 GB or more |
A good rule of thumb: look at the file size and add a few extra gigabytes for context and your operating system. If the total exceeds your available memory, the model will still run, but much more slowly because part of the workload shifts to the processor or the disk.
Installation and your first model #
Download the software from the official LM Studio website for your operating system and install it like any other app.
Upon first launch, it will suggest a model to download. You can accept this or search for a specific one using the magnifying glass in the sidebar.
In the results, each variant shows its file size and whether it is compatible with your hardware. Pick one that fits within your available memory.
Once the download is finished, open the chat, load the model from the top menu, and start typing.
The initial loading takes a few dozen seconds while the model is read into memory. After that, responses are nearly instant.

Essential settings #
When you load a model, you’ll see several parameters. There are three main ones to adjust.
Context Length. This determines how much text the model can keep in its “mind.” Increasing it allows you to work with longer documents, but it consumes significantly more memory; this is the parameter most likely to cause loading errors.
GPU Offloading. This decides how much of the model is loaded onto your graphics card. With a dedicated GPU, set this as high as your VRAM allows: responses will become much faster.
Temperature. This controls how predictable the responses are: use low values for precise tasks like data extraction or classification, and higher values for creative writing.
Using it with your own documents #
The most useful feature in practice is asking questions about your files: a contract, a manual, or a collection of notes. Just drag the PDF into the chat; the program extracts the text, splits it up, and feeds the necessary pieces to the model to answer your question (this technique is called RAG).

This works well for documents of normal size and precise questions. For massive archives or queries that require connecting information scattered across distant points, you will hit limits.
However, for certain uses, the advantage is decisive: nothing leaves your computer, so you can analyze sensitive documents that you would never upload to an external service.
The local server #
LM Studio can run the model as a service on your computer, compatible with the OpenAI API, usually at address http://localhost:1234/v1. This means programs and scripts written for online services can use your local model just by changing the URL. You can also manage it from the terminal using the lms command.

It’s especially interesting for developers: you can try out an application without consuming credits and without sending your data off-device.
Alternatives #
Ollama does the same job via the command line (ollama run qwen3), making it better suited for those who work in the terminal or want to integrate it into scripts and services. Today, it also has its own chat app and pairs well with interfaces like Open WebUI.
Msty, which was the subject of the first version of this article, is now called Msty Studio: it features a very polished interface, lets you manage local models and online services using your API keys in the same window, and includes Knowledge Stacks to query your documents and split chats to compare responses from multiple models. You can start for free, with a paid license available for advanced features.

Integrated system solutions, such as Apple Intelligence models on Apple devices or those on Copilot+ PCs with a dedicated accelerator, give you less choice but require no installation.
GPT4All and similar projects offer an experience close to LM Studio, but with different catalogs and settings.
FAQ #
Does it work without an internet connection? #
Yes, once the model is downloaded. A connection is only needed for the catalog and updates. This is exactly why it’s so useful for anyone working with confidential documents.
How much worse are they than online services? #
For specific tasks (summarizing, rephrasing, translating, or extracting information from text), a good 7-14 billion parameter model performs quite well. However, when it comes to long reasoning, questions requiring extensive knowledge, and complex code, the gap between them and large-scale service models remains significant.
Do I need a dedicated graphics card? #
No, but it helps a lot. Without it, the model runs on your processor and responses arrive more slowly, though small models remain usable. On Macs with Apple silicon, this isn’t an issue because unified memory handles that workload.
Are the models free? #
Open-license ones are, which accounts for almost everything in the catalog. However, licenses vary from model to model: some (like Llama) have restrictions on commercial use, while others (Qwen, Mistral, gpt-oss) are Apache 2.0. If you need a model for work, be sure to check its license.
How much disk space is required? #
Each model takes up anywhere from a couple of gigabytes to several dozen. With four or five models, you’ll quickly hit 50 GB: every once in a while, open the My Models section and remove the ones you no longer use.

