AI Industry Ecosystem
When you choose an AI tool, you are actually choosing a technology path.
Choose closed source or open source? Choose cloud service or local deployment? Choose US companies or Chinese companies? These choices will affect your work efficiency, data security, and long-term costs.
More importantly, if you are considering career development, you need to know which companies are leading the technology and which directions represent the future.
This article will not stuff you with boring industry news, but rather help you build a framework for understanding the AI industry.
Understanding the industry ecosystem is not for joining the hype, but for having confidence when making decisions.
Overview of Major AI Companies
Players in the AI industry can be divided into several categories: leading US companies, leading Chinese companies, and unicorn companies.
Each category of companies has different technology paths and product strategies.
Leading US Companies
The US holds a leading position in large model technology, with four companies being the most representative.
| Company | Core Model | Characteristics | Representative Products |
|---|---|---|---|
| OpenAI | GPT-4、GPT-4o | Technologically leading, mature commercialization | ChatGPT、GPTs |
| Anthropic | Claude 3、Claude 3.5 | Safety-first, strong long-text capabilities | Claude.ai |
| Google DeepMind | Gemini | Strong in multimodality, solid research strength | Gemini App |
| Meta AI | Llama Series | Main force of open source, freely available | Llama 3、Llama 3.1 |
OpenAIIt is the developer of ChatGPT and a key driver in bringing AI to the masses. Its characteristics are fast technology updates and good product experience, but data privacy and pricing strategies often spark controversy.
AnthropicFounded by former OpenAI employees, it focuses on "safe AI". The Claude model performs excellently in handling long documents (tens of thousands of characters) and has relatively lenient content moderation.
Google DeepMindIt is the creator of AlphaGo, now focusing on the Gemini series of models. Its advantages are multimodal capabilities (image, video, audio) and deep research accumulation.
Meta AIIt takes a completely different path—open source. The Llama series models are free for developers to use, and anyone can download, modify, and deploy them. This gives Meta enormous influence in the open source community.
Leading Chinese Companies
Chinese AI companies are also catching up quickly, forming their own ecosystem.
| Company | Core Model | Characteristics | Representative Products |
|---|---|---|---|
| Baidu | Wenxin Large Model Series (ERNIE) | The earliest domestic large model deployment, deeply engaged in multimodality, search, and enterprise services, with rich government and enterprise implementation cases | Wenxin Yiyan, Wenxin Qianfan, Wenxin Yige |
| Alibaba | Tongyi Qianwen Series (Qwen) | Leading open-source efforts, with a full range of lightweight and ultra-large-parameter models, deeply integrated with the e-commerce and cloud computing ecosystem. | Tongyi Qianwen, Tongyi Wanxiang, Alibaba Cloud Bailian |
| Tencent | Hunyuan large model | Leveraging social networking, gaming, and video businesses, it excels in image-text, video generation, digital humans, and consumer-facing social integration scenarios. | Tencent Hunyuan Assistant, Hunyuan Painting, Tencent Zhiying |
| ByteDance | Doubao large model | Outstanding advantages in multimodal and short-video content generation, covering personal conversation, AI creation, and office assistance. | Doubao AI, Dreamina, Coze (Kouzi) |
| Moonshot AI | Moonshot series large models | Focuses on ultra-long context windows, supporting lossless reading of million-level text, with clear advantages in document and knowledge base scenarios. | Kimi intelligent assistant |
| Zhipu AI | GLM (Zhipu Qingyan) | Strong academic background, covering general dialogue, code, and scientific research scenarios, with mature enterprise private deployment. | Zhipu Qingyan, Zhipu Code, Zhipu Qingyan Enterprise Edition |
| MiniMax (Xiyu Technology) | 01 large model | Outstanding capabilities in voice, digital humans, and video generation, with a well-established overseas commercialization layout. | Hailuo AI, digital human video generation platform |
| DeepSeek | DeepSeek large model | Top-tier code large model performance, large-scale open-sourcing of the entire model series, targeting developers and researchers. | DeepSeek Chat、DeepSeek Coder |
BaiduIt is one of the earliest companies in China to invest in large models. ERNIE Bot has an advantage in Chinese understanding.
AlibabaThe Tongyi Qianwen (Qwen) series models are very actively open-sourced, ranging from 1.8 billion to 72 billion parameters, with high usage in the open-source community.
Dark Side of the Moon (Moonshot)It is a startup focusing on ultra-long context—its Kimi assistant can process documents of over one million Chinese characters at a time, which is very useful for handling books and codebases.
Unicorn Companies
There are also some startups that have not been around long, but have distinctive technological features and high valuations.
| Company | Core models | Characteristics | Headquarters |
|---|---|---|---|
| Mistral AI | Mistral series | Technological innovation, efficiency first | France |
| Cohere | Command R+ | Enterprise services, retrieval-augmented | Canada |
| xAI | Grok | Founded by Musk, pursuing truth | United States |
Mistral AIIt is a French company, known for its compact yet powerful models. Mistral 7B leads in performance at the same parameter scale, and it has introduced the MoE (Mixture of Experts) architecture, which is highly efficient.
Which company's product to choose mainly depends on your needs: for the strongest capability, choose OpenAI or Anthropic; for free open source, choose DeepSeek, Zhipu, etc.
Open Source vs Closed Source Models
This is the most important strategic debate in the AI industry.
On one side is the closed-source route represented by OpenAI and Anthropic—models are only accessible via API or web, and you cannot see the model weights.
On the other side is the open-source route represented by DeepSeek and Alibaba—you can download complete models and run them on your own devices.
What are Open Source Models?
Open source has a clear definition in the software field, but in the AI field, the situation is somewhat different.
True open source means you can freely:
- Download the model weights
- Use for commercial purposes
- Modify the model architecture
- Republish derivative versions
But many so-called open-source models are actually open-weight—you can download and use them, but there are commercial use restrictions.
Typical Open Source Models
| Model series | Developer | License | Characteristics |
|---|---|---|---|
| Llama 3/3.1 | Meta | Business-friendly | Most complete ecosystem, good community support |
| Mistral | Mistral AI | Apache 2.0 | High efficiency, many innovations |
| Qwen | Alibaba | Apache 2.0 | Strong Chinese language capability |
| Yi | 01.AI | Business-friendly | Balanced in Chinese and English |
Closed Source vs Open Source Comparison
The two routes each have strengths and weaknesses; there is no absolute good or bad—it depends on your scenario.
| Dimension | Closed-source models | Open-source models |
|---|---|---|
| Capability ceiling | Usually stronger (GPT-4, Claude 3) | Catching up, but the gap is narrowing |
| Data privacy | Data sent to third parties | Data stays local |
| Usage cost | Pay per token | One-time hardware investment |
| Customization capability | Limited (can only fine-tune prompts) | Fully controllable (fine-tuning, quantization) |
| Deployment difficulty | Zero deployment (use API directly) | Requires technical skills to deploy |
| Update speed | Vendor continuously updates | Need to track new versions yourself |
How to Choose
Choose closed-source when:
- You want the strongest capabilities and don't want to deal with technical details
- Your data is not sensitive and can be sent to third parties
- Your usage is small, pay-per-token is more cost-effective
- You need to launch quickly and don't want to spend time on deployment
Choose open-source when:
- Your data is highly sensitive and cannot leave the company
- Your usage is large, pay-per-token is too expensive
- You need deep customization of model behavior
- You have a technical team to deploy and maintain
Recommended strategy: use closed-source for small usage and exploration (low cost, fast launch), consider open-source for large-scale and production environments (strong controllability, low long-term cost).
Cloud AI vs Local AI
This is another important choice: use cloud AI services, or run models on your own computer?
Cloud AI: Ready to Use
Cloud AI means calling AI services in the cloud via web pages or APIs.
You don't need to worry about where the model is deployed, how many GPUs are used, or how it's updated—you just input content and get results.
Typical cloud AI products:
| Product | Type | Payment method |
|---|---|---|
| ChatGPT | Web chat | Subscription $20/month, or pay-as-you-go |
| Claude.ai | Web chat | Subscription $20/month, or pay-as-you-go |
| OpenAI API | API interface | Billed by token |
| Anthropic API | API interface | Billed by token |
Local AI: Fully Controllable
Local AI runs open-source models on your own computer or server.
Data doesn't leave your device, and there's no ongoing payment.
But you need a good graphics card—usually an NVIDIA GPU; the more VRAM the better.
Common tools for local AI:
| Tool | Features | Suitable for |
|---|---|---|
| Ollama | One-click start, command-line operation | Developers, technical users |
| LM Studio | GUI, rich model library | General users, beginners |
| vLLM | High-performance inference service | Production environment deployment |
| Text Generation WebUI | Most feature-complete, fine-tunable | Advanced users |
Cloud AI vs Local AI Comparison
| Dimension | Cloud AI | Local AI |
|---|---|---|
| Initial cost | Zero, pay as you go | High (need to buy GPU) |
| Usage cost | Continuous pay-per-use | Electricity cost + hardware depreciation |
| Data privacy | Data sent to service provider | Data stays entirely local |
| Deployment difficulty | None | Requires some technical skills |
| Response speed | Depends on network | Local computing, stable speed |
| Capability ceiling | Can use the largest models | Limited by GPU VRAM |
Hybrid Strategy: Getting the Best of Both Worlds
Many companies adopt a hybrid strategy:
- Sensitive data uses local models
- Ordinary queries use cloud APIs
- Use cloud services during daytime peaks, local for nighttime batch processing
This ensures data security while controlling costs.
Advice for individual users: try cloud AI first, and after getting familiar, consider local deployment if you have privacy needs or usage increases.
AI Infrastructure
AI is not just models—the underlying computing power, chips, and cloud services form the infrastructure of the entire industry.
GPU Computing Power: Why NVIDIA Dominates the Market
Training large models requires massive computation, and GPUs (Graphics Processing Units) are currently the most suitable hardware.
NVIDIA holds a near-monopoly in this field, for three reasons:
| Advantage | Description |
|---|---|
| CUDA ecosystem | CUDA is NVIDIA's programming framework; almost all AI frameworks prioritize support for it. |
| Memory capacity | Professional cards like H100 and A100 have 80GB or even more memory. |
| Software support | Optimized libraries, compilers, and toolchains are the most complete. |
Common NVIDIA GPU positioning:
| Card model | Memory | Positioning | Applicable scenarios |
|---|---|---|---|
| RTX 3090/4090 | 24GB | Consumer-level | Personal learning, small model inference |
| A10G | 24GB | Entry-level professional | Lightweight inference, fine-tuning |
| A100 | 80GB | Professional-level | Model training, large-scale inference |
| H100 | 80GB | Flagship-level | Large model training (70B+) |
Memory size is a key metric.Generally speaking:
- A 7B model requires 8-16GB of memory
- A 13B model requires 16-32GB of memory
- A 70B model requires 60-80GB of memory
Cloud Computing Providers: Compute as a Service
Not everyone can afford a GPU costing hundreds of thousands—cloud vendors offer a "pay-per-hour rental" model.
| Cloud vendor | Features | Typical instances |
|---|---|---|
| AWS | Most complete instance types | p4d, p5, g5 series |
| Azure | Complete enterprise services | NC, ND, NV series |
| GCP | Good TPU support | A3, T4 series |
| Alibaba Cloud | Fast access in China | gn6, gn7 series |
Domestic Computing Power Alternatives
Due to chip export restrictions, China is actively developing domestic computing power.
Main players:
| Vendor | Chip | Features |
|---|---|---|
| Huawei Ascend | Ascend 910 | Ecosystem relatively complete, supports large model training |
| Cambricon | Siyuan 590 | Focused on AI acceleration |
| Hygon | DCU | CUDA-ecosystem compatible |
Domestic chips are still catching up; the main issue is the software ecosystem—many AI frameworks require special adaptation.
For individual developers, it's recommended to start with NVIDIA consumer-grade GPUs (e.g., RTX 4090), as the ecosystem is the most mature and has the most resources.
AI Business Models
How do AI companies make money? Currently there are three main models.
B2C: Subscription Model
Targets individual users, paid monthly.
| Product | Pricing | What's included |
|---|---|---|
| ChatGPT Plus | $20/month | Priority access to GPT-4, faster responses |
| Claude Pro | $20/month | Priority access to Claude 3, longer context |
| GitHub Copilot | $10/month | Code completion, code explanation |
The benefit of the subscription model is stable cash flow and high user loyalty. However, it must continuously deliver value, otherwise users will cancel.
B2B: API Calls
Pay-as-you-go—you pay based on how many times you call the API.
The pricing unit is "Token," which can be understood as "word" or "character."
Typical pricing (mid-2025):
| Model | Input price | Output price |
|---|---|---|
| GPT-4o | $5 per million tokens | $15 per million tokens |
| Claude 3.5 Sonnet | $3 per million tokens | $15 per million tokens |
| GPT-4 | $30 per million tokens | $60 per million tokens |
| Llama 3.1 405B (self-hosted) | Hardware cost + electricity | |
Let's do the math: Suppose you have an app that processes 100,000 queries per day, with each query using 1,000 tokens.
With GPT-4o, the daily cost is $5 (input) + $15 (output) = $20, which comes to $600 per month.
If usage grows, costs rise quickly — which is why many companies ultimately choose to self-host open-source models.
Monetizing the Open Source Ecosystem
Open-source models are free, but companies can make money through peripheral services:
- Hosting services: deploy and maintain for you
- Enterprise edition: provide extra features and support
- Consulting services: customize and optimize for you
- Training services: fine-tune models with your data
Mistral AI is a representative of this model — the models are open source, but enterprises need to pay to access larger models and better services.
Comparison of the Three Models
| Model | Target users | Revenue stability | Technical requirements |
|---|---|---|---|
| Subscription | Individual users | Medium | Requires excellent product experience |
| API calls | Developers, enterprises | High (usage-driven) | Requires stable service |
| Open-source ecosystem | Enterprises, developers | Slower (requires building an ecosystem) | Requires community operations capability |
AI Regulation and Policy
AI is developing so fast that every country is formulating regulatory policies. Understanding these policies can help you mitigate risks.
EU AI Act
The EU has the strictest regulation, with the core concept being "risk-based classification."
| Risk level | Definition | Regulatory requirements |
|---|---|---|
| Unacceptable risk | Social scoring, real-time facial recognition | Prohibited |
| High risk | Medical devices, legal judgments, recruitment | Strict compliance, high transparency |
| Medium risk | Chatbots, content generation | Label AI-generated content |
| Low risk | Gaming, photo editing | Essentially no requirements |
The EU AI Act has a global impact — as long as your product serves EU users, you must comply.
China's AI Regulatory Framework
China's regulatory policies are also being gradually refined, with core documents including:
- "Interim Measures for the Management of Generative AI Services"
- "Artificial Intelligence Law" (Draft for Comments)
Key requirements:
- Generated content must be truthful and accurate, and must not fabricate false information
- Content review mechanisms must be sound
- Data sources must be legitimate
- Algorithms must be transparent and explainable
US AI Policy Dynamics
US regulation is relatively lenient, primarily through executive orders and industry self-regulation.
Key focus areas:
- National security: prevent AI from being used for malicious purposes
- Fairness: prevent algorithmic discrimination
- Transparency: AI-generated content needs to be labeled
For individual developers, the most important things to note are: AI-generated content should be clearly labeled and should not masquerade as human creation. When personal data is involved, data protection regulations must be followed.
2025 AI Industry Trend Outlook
As of mid-2026, several trends have become quite clear.
Trend 1: Model Capabilities Continue to Improve, but with Diminishing Marginal Returns
GPT-5, Claude 5, Gemini 3.0, DeepSeek 4, and GLM 5 have all arrived, but each upgrade brings diminishing surprises.
It's not that technology has hit a bottleneck, but rather that there are already many "good enough" use cases.
Trend 2: Open Source Models Continue to Catch Up, Narrowing the Gap with Closed Source
Llama 3.1 405B's performance is already close to GPT-4's level, and it's freely available.
More companies will join the open-source camp in the future, and the open-source model ecosystem will become increasingly mature.
Trend 3: The Application Layer Begins to Explode, Agents Become Practical
The past few years were about competition at the model layer; now the focus is shifting to the application layer.
AI Agents — AI that can autonomously complete complex tasks — will move from concept to practical application.
Trend 4: Multimodality Becomes Standard
Pure text models will become increasingly rare, and multimodal models that can handle text, images, audio, and video simultaneously will become the standard.
Trend 5: Regulation Becomes Normalized
Regulatory frameworks in various countries will gradually take shape, moving the AI industry from "reckless growth" into a stage of "regulated development."
| Trends | Impact on individuals | Recommendations |
|---|---|---|
| Model capability improvement | Tools are becoming more powerful | Watch for new features, but don't blindly chase novelty |
| Open-source models catching up | Local deployment is becoming more feasible | Learn the open-source tool stack |
| Explosion at the application layer | Opportunities lie at the application layer | Combine AI with your domain |
| Multimodal as standard | Can do more things | Expand multimodal application ideas |
| Regulatory normalization | Compliance requirements are increasing | Understand relevant regulations |
Code Example: Calling an Open Source Model API
Let's write a simple example in Python to demonstrate how to call an open-source model API.
We'll use Ollama as the local inference service, which supports multiple open-source models such as Llama 3, Qwen, and Mistral.
Reference for Ollama-related content:https://www.example.com/ollama/ollama-tutorial.html。
Example
# Example of calling an open-source model API
# Demo: Running the Llama 3 model locally with Ollama
# ============================================
import requests
import json
def call_ollama_model(prompt: str, model: str = "llama3") -> str:
"""
Call Ollama's API to get a model reply
Parameter descriptions:
prompt: Your question or input text
model: The model name to use (default: llama3)
Optional values: llama3, qwen2, mistral, yi, etc.
Return value:
The model's reply text
"""
# Ollama runs on local port 11434 by default
url = "http://localhost:11434/api/generate"
# Build the request data
payload = {
"model": model, # Model to use
"prompt": prompt, # User input
"stream": False, # Do not use streaming output (return at once)
"options": {
"temperature": 0.7, # Temperature parameter: higher is more random, lower is more deterministic
"top_p": 0.9, # Sampling parameter: controls diversity
"num_predict": 512, # Maximum number of generated tokens
}
}
try:
# Send POST request
response = requests.post(url, json=payload, timeout=60)
response.raise_for_status() # Check whether the request succeeded
# Parse the returned result
result = response.json()
return result.get("response", "No reply obtained")
except requests.exceptions.ConnectionError:
return "Error: Unable to connect to the Ollama service. Please start Ollama first."
except requests.exceptions.Timeout:
return "Error: Request timed out. The model may still be loading."
except Exception as e:
return f"Error: {str(e)}"
def chat_with_example():
"""Example of multi-turn conversation with the example AI assistant"""
print("=" * 50)
print("🚀 example AI assistant (based on Llama 3 open-source model)")
print("=" * 50)
print("Tip: Enter 'quit' or 'exit' to exit the conversation")
print()
# Conversation history (for context continuity)
conversation_history = []
while True:
# Get user input
user_input = input("You: ").strip()
# Check whether to exit
if user_input.lower() in ["quit", "exit", "Exiting"]:
print("example AI: Goodbye! Welcome back anytime.")
break
if not user_input:
continue
# Build a prompt with context
full_prompt = "\n".join(conversation_history) + f"\nUser: {user_input}\nAssistant:"
# Call the model
print("example AI: ", end="", flush=True)
response = call_ollama_model(full_prompt)
print(response)
print()
# Update conversation history
conversation_history.append(f"User: {user_input}")
conversation_history.append(f"Assistant: {response}")
# Keep only the last 5 turns of conversation to avoid overly long context
if len(conversation_history) > 10:
conversation_history = conversation_history[-10:]
# Example 1: single call
print("--- Example 1: single question ---")
question = "Please explain what a large language model is in one sentence."
print(f"Question: {question}")
answer = call_ollama_model(question)
print(f"Answer: {answer}")
print()
# Example 2: using different models
print("--- Example 2: switching between different models ---")
for model_name in ["llama3", "qwen2", "mistral"]:
print(f"Using model {model_name}:")
answer = call_ollama_model("Hello, please introduce yourself.", model=model_name)
print(f"Answer: {answer[:100]}...") # Only display the first 100 characters
print()
# Example 3: multi-turn conversation (uncomment to run)
# print("--- Example 3: multi-turn conversation ---")
# chat_with_example()
Before running this code, you need to install Ollama first:
- gohttps://ollama.comDownload and install Ollama
- Run in terminal:
ollama pull llama3Download the model - Run:
ollama serveStart the service (if not already started) - Then run the Python code above
If you don't have a local GPU, you can also use cloud-based open-source model APIs, such as Together.ai, Anyscale, etc. The calling method is similar, except the URL and authentication method are different.
Other extensionsThe advantage of open-source models is that you can first use a small model (7B) locally for development and testing, and if it works well, switch to a larger model (70B) or cloud service.