↓Skip to main content
Alemasone

Run AI Models Locally with LM Studio

Learn how to install LM Studio, choose the right model for your hardware, and use offline AI: requirements, quantization, and getting started.

7 min read Updated on
LM Studio interface with a loaded model and an ongoing conversation
On this page
If you run a language model on your computer, your documents never leave your machine and you don’t pay any subscription fees. LM Studio is the easiest way to do it: just install it, choose a model from the catalog, and start typing. The limiting factor is memory: you need at least as much RAM as the size of the model file, plus some extra headroom.

Setting realistic expectations. Models running on a home computer are less capable than those used by major online services. For summarizing, paraphrasing, classifying, translating, and reasoning over your documents, they work great; however, for complex tasks, you will notice the difference compared to ChatGPT or Claude.

As of July 2025, LM Studio is also free for business use, with no licenses required.

Requirements #

SystemRequirements
macOSApple Silicon only, macOS 14 or later; 16 GB RAM recommended
Windowsx64 with AVX2 instructions or ARM; 16 GB RAM and 4 GB VRAM recommended
LinuxUbuntu 20.04 or later, distributed as an AppImage for x64 and ARM64

It works with 8 GB of memory, but only with the smallest models. The real jump happens at 16 GB, and with 32 GB, you can run almost anything that makes sense on a personal computer.

On Mac with Apple silicon memory is shared between the processor and graphics, which is a huge advantage: a Mac with 32 GB can load models that would require a high-end graphics card on a PC. LM Studio on Mac also uses models in MLX format, which are optimized for those chips and are usually faster. On a PC with an NVIDIA card, make sure to keep your drivers up to date.

How much memory do you need for each model #

That is the key question, and the answer depends on quantization: a technique that reduces the precision of the numbers within the model, making the file much lighter at the cost of a slight drop in quality.

A model is identified by its number of parameters (7B, 14B, 32B, where B stands for billions) and its quantization level, often written as Q4_K_M or similar. The lower the number after the Q, the lighter the file and the more the quality decreases.

For practical reference:

Model4-bit quantized fileRecommended memory
3Bapprox. 2 GB8 GB
7-8Bapprox. 4-5 GB16 GB
14Bapprox. 8-9 GB16-24 GB
32Bapprox. 18-20 GB32 GB
70Bapprox. 40 GB64 GB or more

A good rule of thumb: look at the file size and add a few extra gigabytes for context and your operating system. If the total exceeds your available memory, the model will still run, but much more slowly because part of the workload shifts to the processor or the disk.

Where to start. Choose a 7-8 billion parameter model quantized at 4 bits—for example, Qwen 3 8B, Llama 3.1 8B, or Gemma 3 12B if you have 16 GB. These run on almost anything, respond at a reasonable speed, and are enough to help you decide if this workflow works for you. With 16 GB of VRAM or a 32 GB Mac, you can even try gpt-oss 20B, OpenAI’s open model.

Installation and your first model #

  1. Download the software from the official LM Studio website for your operating system and install it like any other app.

  2. Upon first launch, it will suggest a model to download. You can accept this or search for a specific one using the magnifying glass in the sidebar.

  3. In the results, each variant shows its file size and whether it is compatible with your hardware. Pick one that fits within your available memory.

  4. Once the download is finished, open the chat, load the model from the top menu, and start typing.

The initial loading takes a few dozen seconds while the model is read into memory. After that, responses are nearly instant.

Searching for models in LM Studio using the model list and quantized variant download options
The search indicates which variants fit within your available memory

Essential settings #

When you load a model, you’ll see several parameters. There are three main ones to adjust.

Context Length. This determines how much text the model can keep in its “mind.” Increasing it allows you to work with longer documents, but it consumes significantly more memory; this is the parameter most likely to cause loading errors.

GPU Offloading. This decides how much of the model is loaded onto your graphics card. With a dedicated GPU, set this as high as your VRAM allows: responses will become much faster.

Temperature. This controls how predictable the responses are: use low values for precise tasks like data extraction or classification, and higher values for creative writing.

Using it with your own documents #

The most useful feature in practice is asking questions about your files: a contract, a manual, or a collection of notes. Just drag the PDF into the chat; the program extracts the text, splits it up, and feeds the necessary pieces to the model to answer your question (this technique is called RAG).

LM Studio chat with an attached PDF and the model’s response based on its content

This works well for documents of normal size and precise questions. For massive archives or queries that require connecting information scattered across distant points, you will hit limits.

However, for certain uses, the advantage is decisive: nothing leaves your computer, so you can analyze sensitive documents that you would never upload to an external service.

The local server #

LM Studio can run the model as a service on your computer, compatible with the OpenAI API, usually at address http://localhost:1234/v1. This means programs and scripts written for online services can use your local model just by changing the URL. You can also manage it from the terminal using the lms command.

LM Studio local server settings showing port 1234, CORS, and loading models on demand

It’s especially interesting for developers: you can try out an application without consuming credits and without sending your data off-device.

Alternatives #

Ollama does the same job via the command line (ollama run qwen3), making it better suited for those who work in the terminal or want to integrate it into scripts and services. Today, it also has its own chat app and pairs well with interfaces like Open WebUI.

Msty, which was the subject of the first version of this article, is now called Msty Studio: it features a very polished interface, lets you manage local models and online services using your API keys in the same window, and includes Knowledge Stacks to query your documents and split chats to compare responses from multiple models. You can start for free, with a paid license available for advanced features.

Msty Chat with model selection and linked knowledge stacks
Msty combines local models, online services, and documents

Integrated system solutions, such as Apple Intelligence models on Apple devices or those on Copilot+ PCs with a dedicated accelerator, give you less choice but require no installation.

GPT4All and similar projects offer an experience close to LM Studio, but with different catalogs and settings.

FAQ #

Does it work without an internet connection? #

Yes, once the model is downloaded. A connection is only needed for the catalog and updates. This is exactly why it’s so useful for anyone working with confidential documents.

How much worse are they than online services? #

For specific tasks (summarizing, rephrasing, translating, or extracting information from text), a good 7-14 billion parameter model performs quite well. However, when it comes to long reasoning, questions requiring extensive knowledge, and complex code, the gap between them and large-scale service models remains significant.

Do I need a dedicated graphics card? #

No, but it helps a lot. Without it, the model runs on your processor and responses arrive more slowly, though small models remain usable. On Macs with Apple silicon, this isn’t an issue because unified memory handles that workload.

Are the models free? #

Open-license ones are, which accounts for almost everything in the catalog. However, licenses vary from model to model: some (like Llama) have restrictions on commercial use, while others (Qwen, Mistral, gpt-oss) are Apache 2.0. If you need a model for work, be sure to check its license.

How much disk space is required? #

Each model takes up anywhere from a couple of gigabytes to several dozen. With four or five models, you’ll quickly hit 50 GB: every once in a while, open the My Models section and remove the ones you no longer use.

Read next