Falcon
FreeOpen LLM family from UAE's Technology Innovation Institute with the permissive TII Falcon License
What is Falcon?
Falcon is the open-source LLM family developed by the Technology Innovation Institute (TII) in Abu Dhabi, first launched in 2023 with Falcon 7B and 40B and now several generations past the Falcon 3 release. Falcon earned its reputation early by briefly topping the Hugging Face Open LLM Leaderboard with Falcon 40B, and TII kept pushing with Falcon 180B and Falcon 2 multimodal in 2024, then Falcon 3 and the state-space Falcon Mamba 7B. The current line-up is anchored by Falcon-H1, a hybrid Transformer plus Mamba architecture shipped in 0.5B, 1.5B, 1.5B-Deep, 3B, 7B, and 34B base and instruct variants, alongside Falcon-E for edge hardware where it runs on CPUs rather than GPUs, Falcon H1R 7B for reasoning, Falcon-H1-Arabic, the very small Falcon-H1-Tiny and Tiny-R checkpoints, and Falcon Perception, a multimodal model that reads and interprets images from natural-language prompts. TII claims Falcon-H1 models outperform models twice their size, with the 0.5B checkpoint landing close to typical 7B models. Falcon models are released under the TII Falcon License, in part based on Apache 2.0 with additional acceptable-use provisions, in practical terms, Falcon is free to download, fine-tune, and use commercially for the vast majority of applications. The widely repeated managed-hosting restriction, which requires a separate TII agreement before offering shared instances as an inference or fine-tuning API, actually lives in the older Falcon 180B license and is not carried into the December 2024 TII Falcon License that covers Falcon 3 and the Falcon-H1 generation. Where Falcon stands out is as a credible non-US, non-Chinese open LLM option: for governments, enterprises, and researchers who want to diversify their LLM supply chain away from Meta, Google, and Alibaba, Falcon offers an independent pretraining lineage with different data sources, languages, and safety tuning. It is particularly strong in Arabic and MENA-region languages where major Western models lag.
Falcon demo video
Watch ATRC's official demo to see Falcon in action before reading our full review.
Official video by ATRC via YouTube, embedded for reference. ToolChase does not host or claim this video.
⚡ Quick Verdict
Teams that want a non-US/non-China open LLM or need strong Arabic language capabilities
Users who want the largest ecosystem or zero-fine-print Apache 2.0 licensing
Free to download · Inference via third parties
Yes, weights are free under the TII Falcon License
Independent pretraining lineage with strong Arabic support and efficient hybrid Falcon-H1 models
Smaller ecosystem and less permissive license than Gemma
Bottom line: Falcon scores 4.2/5, the top pick when geographic LLM diversification or Arabic language support matters. Choose Falcon-E or a small Falcon-H1 checkpoint for laptops and edge, Falcon-H1 34B for the strongest current open Falcon.
Pricing
Weights, Free: All Falcon models can be downloaded from Hugging Face or falconllm.tii.ae under the TII Falcon License. Commercial use is permitted for the vast majority of applications without fees.
Inference providers: Falcon is available on Hugging Face Inference, Amazon Bedrock Marketplace, AWS SageMaker JumpStart, and community-hosted endpoints at standard per-token rates.
Self-hosting: Falcon-E is designed to run on CPUs rather than GPUs, and the small Falcon-H1 checkpoints (0.5B to 3B) plus Falcon-H1-Tiny run comfortably on a laptop. Falcon-H1 7B suits a single consumer GPU; Falcon-H1 34B and Falcon 180B need server-class multi-GPU setups.
Licensing note: The requirement to negotiate a separate TII agreement before offering shared managed instances as an inference or fine-tuning API sits in the Falcon 180B license. The December 2024 TII Falcon License, which covers Falcon 3 and the Falcon-H1 generation, does not include that clause, and Falcon 7B and Falcon 40B remain under plain Apache 2.0.
Key Features
- TII Falcon License, in part based on Apache 2.0, permissive for most commercial use
- Falcon-H1 hybrid Transformer plus Mamba generation: 0.5B, 1.5B, 1.5B-Deep, 3B, 7B, 34B
- Falcon H1R 7B for math, coding, logic, and instruction-following reasoning
- Falcon-E for edge deployment, built to run on CPUs rather than GPUs
- Falcon-H1-Tiny and Tiny-R down to 0.09B for tiny-footprint reasoning
- Falcon Perception, a multimodal model that reads and understands images
- Falcon-H1-Arabic and Falcon Arabic for MENA-region language support
- Earlier generations still published: Falcon 3, Falcon Mamba 7B, Falcon 2, Falcon 180B, Falcon 40B
- Pre-trained on RefinedWeb dataset
- Available on Hugging Face, Amazon Bedrock Marketplace, and AWS SageMaker JumpStart
- Backed by UAE's Technology Innovation Institute
Pros & Cons
Pros
- Independent non-US/non-China open LLM option
- Excellent Arabic and MENA language performance
- Efficient Falcon-H1 and Falcon-E small models for edge and CPU deployment
- Permissive license for most commercial use cases
Cons
- Smaller community and ecosystem than Llama or Mistral
- License has more fine print than Apache 2.0
- Falcon 180B's managed-hosting clause can complicate shared-inference SaaS deployments
FAQ
Is Falcon free for commercial use?
For almost everyone, yes. Falcon models released under the TII Falcon License, which is in part based on Apache 2.0, can be used commercially, modified, and redistributed without fees. The hosting caveat is narrower than it is usually reported: the clause requiring a separate TII agreement before offering shared managed instances as an inference or fine-tuning API sits in the older Falcon 180B license, and the December 2024 TII Falcon License that covers Falcon 3 and the current Falcon-H1 generation does not carry it. Falcon 7B and Falcon 40B are plain Apache 2.0. For internal enterprise use, building applications on top of Falcon, and most SaaS scenarios, no license negotiation is needed.
Which Falcon model should I use?
For laptop and edge deployment, Falcon-E is built to run on CPUs rather than GPUs, and the small Falcon-H1 checkpoints at 0.5B, 1.5B, and 3B are the most efficient general options, with Falcon-H1-Tiny going smaller still. For server inference on a single consumer GPU, Falcon-H1 7B hits the best capability/cost balance, and Falcon H1R 7B is the pick when the workload is math, coding, logic, or instruction-following. Falcon-H1 34B is the strongest current general-purpose Falcon. Falcon 180B is now a legacy option: it still offers heavyweight open performance but costs far more to run and carries the stricter hosting license.
How does Falcon compare to Llama?
Meta Llama 4 and Llama 3.3 70B outperform Falcon on most English-language benchmarks and have a much larger fine-tune ecosystem. Falcon wins in three scenarios: geographic diversification away from US model providers, Arabic and MENA-region language tasks, and efficient Falcon-H1 and Falcon-E small models for edge deployment.
Is Falcon good for Arabic?
Yes, Arabic and MENA-region language support is one of Falcon's biggest differentiators. TII is UAE-based, and the pretraining data includes significantly more high-quality Arabic text than most Western open models. TII released Falcon-H1-Arabic in 3B, 7B, and 34B sizes in January 2026 and reports it as the top performer on the Open Arabic LLM Leaderboard. For Arabic-first applications, Falcon generally outperforms Llama, Gemma, and Mistral on MENA benchmarks.
Where can I run Falcon?
You can download weights from Hugging Face and run locally with Transformers, Ollama, or vLLM. Managed hosting is available on Amazon Bedrock Marketplace and AWS SageMaker JumpStart, and Hugging Face Inference Endpoints can serve Falcon with pay-per-hour GPU pricing.
What is the RefinedWeb dataset?
RefinedWeb is TII's curated pretraining dataset filtered from Common Crawl with aggressive deduplication and quality filters. It was one of the first large-scale demonstrations that high-quality web data alone can match or beat curated corpora for LLM pretraining, and it remains open for research use on Hugging Face.
📋 Good to know
Download from Hugging Face with one line of Transformers code, or use Ollama/llama.cpp for the small Falcon-H1 and Falcon-E models.
Self-hosting keeps all data on your infrastructure. TII does not collect inference telemetry from downloaded weights.
Move from Falcon-H1 7B to Falcon-H1 34B for better general reasoning, or to Falcon H1R 7B when the task is math, code, or logic.
Low, standard Hugging Face Transformers loading with the tiiuae organization prefix.