AI Industry Ecosystem

When you choose an AI tool, you are actually choosing a technology path.

Choose closed source or open source? Choose cloud service or local deployment? Choose US companies or Chinese companies? These choices will affect your work efficiency, data security, and long-term costs.

More importantly, if you are considering career development, you need to know which companies are leading the technology and which directions represent the future.

This article will not stuff you with boring industry news, but rather help you build a framework for understanding the AI industry.

Understanding the industry ecosystem is not for joining the hype, but for having confidence when making decisions.


Overview of Major AI Companies

Players in the AI industry can be divided into several categories: leading US companies, leading Chinese companies, and unicorn companies.

Each category of companies has different technology paths and product strategies.

Leading US Companies

The US holds a leading position in large model technology, with four companies being the most representative.

CompanyCore ModelCharacteristicsRepresentative Products
OpenAIGPT-4、GPT-4oTechnologically leading, mature commercializationChatGPT、GPTs
AnthropicClaude 3、Claude 3.5Safety-first, strong long-text capabilitiesClaude.ai
Google DeepMindGeminiStrong in multimodality, solid research strengthGemini App
Meta AILlama SeriesMain force of open source, freely availableLlama 3、Llama 3.1

OpenAIIt is the developer of ChatGPT and a key driver in bringing AI to the masses. Its characteristics are fast technology updates and good product experience, but data privacy and pricing strategies often spark controversy.

AnthropicFounded by former OpenAI employees, it focuses on "safe AI". The Claude model performs excellently in handling long documents (tens of thousands of characters) and has relatively lenient content moderation.

Google DeepMindIt is the creator of AlphaGo, now focusing on the Gemini series of models. Its advantages are multimodal capabilities (image, video, audio) and deep research accumulation.

Meta AIIt takes a completely different path—open source. The Llama series models are free for developers to use, and anyone can download, modify, and deploy them. This gives Meta enormous influence in the open source community.

Leading Chinese Companies

Chinese AI companies are also catching up quickly, forming their own ecosystem.

CompanyCore ModelCharacteristicsRepresentative Products
BaiduWenxin Large Model Series (ERNIE)The earliest domestic large model deployment, deeply engaged in multimodality, search, and enterprise services, with rich government and enterprise implementation casesWenxin Yiyan, Wenxin Qianfan, Wenxin Yige
AlibabaTongyi Qianwen Series (Qwen)Leading open-source efforts, with a full range of lightweight and ultra-large-parameter models, deeply integrated with the e-commerce and cloud computing ecosystem.Tongyi Qianwen, Tongyi Wanxiang, Alibaba Cloud Bailian
TencentHunyuan large modelLeveraging social networking, gaming, and video businesses, it excels in image-text, video generation, digital humans, and consumer-facing social integration scenarios.Tencent Hunyuan Assistant, Hunyuan Painting, Tencent Zhiying
ByteDanceDoubao large modelOutstanding advantages in multimodal and short-video content generation, covering personal conversation, AI creation, and office assistance.Doubao AI, Dreamina, Coze (Kouzi)
Moonshot AIMoonshot series large modelsFocuses on ultra-long context windows, supporting lossless reading of million-level text, with clear advantages in document and knowledge base scenarios.Kimi intelligent assistant
Zhipu AIGLM (Zhipu Qingyan)Strong academic background, covering general dialogue, code, and scientific research scenarios, with mature enterprise private deployment.Zhipu Qingyan, Zhipu Code, Zhipu Qingyan Enterprise Edition
MiniMax (Xiyu Technology)01 large modelOutstanding capabilities in voice, digital humans, and video generation, with a well-established overseas commercialization layout.Hailuo AI, digital human video generation platform
DeepSeekDeepSeek large modelTop-tier code large model performance, large-scale open-sourcing of the entire model series, targeting developers and researchers.DeepSeek Chat、DeepSeek Coder

BaiduIt is one of the earliest companies in China to invest in large models. ERNIE Bot has an advantage in Chinese understanding.

AlibabaThe Tongyi Qianwen (Qwen) series models are very actively open-sourced, ranging from 1.8 billion to 72 billion parameters, with high usage in the open-source community.

Dark Side of the Moon (Moonshot)It is a startup focusing on ultra-long context—its Kimi assistant can process documents of over one million Chinese characters at a time, which is very useful for handling books and codebases.

Unicorn Companies

There are also some startups that have not been around long, but have distinctive technological features and high valuations.

CompanyCore modelsCharacteristicsHeadquarters
Mistral AIMistral seriesTechnological innovation, efficiency firstFrance
CohereCommand R+Enterprise services, retrieval-augmentedCanada
xAIGrokFounded by Musk, pursuing truthUnited States

Mistral AIIt is a French company, known for its compact yet powerful models. Mistral 7B leads in performance at the same parameter scale, and it has introduced the MoE (Mixture of Experts) architecture, which is highly efficient.

Which company's product to choose mainly depends on your needs: for the strongest capability, choose OpenAI or Anthropic; for free open source, choose DeepSeek, Zhipu, etc.


Open Source vs Closed Source Models

This is the most important strategic debate in the AI industry.

On one side is the closed-source route represented by OpenAI and Anthropic—models are only accessible via API or web, and you cannot see the model weights.

On the other side is the open-source route represented by DeepSeek and Alibaba—you can download complete models and run them on your own devices.

What are Open Source Models?

Open source has a clear definition in the software field, but in the AI field, the situation is somewhat different.

True open source means you can freely:

  • Download the model weights
  • Use for commercial purposes
  • Modify the model architecture
  • Republish derivative versions

But many so-called open-source models are actually open-weight—you can download and use them, but there are commercial use restrictions.

Typical Open Source Models

Model seriesDeveloperLicenseCharacteristics
Llama 3/3.1MetaBusiness-friendlyMost complete ecosystem, good community support
MistralMistral AIApache 2.0High efficiency, many innovations
QwenAlibabaApache 2.0Strong Chinese language capability
Yi01.AIBusiness-friendlyBalanced in Chinese and English

Closed Source vs Open Source Comparison

The two routes each have strengths and weaknesses; there is no absolute good or bad—it depends on your scenario.

DimensionClosed-source modelsOpen-source models
Capability ceilingUsually stronger (GPT-4, Claude 3)Catching up, but the gap is narrowing
Data privacyData sent to third partiesData stays local
Usage costPay per tokenOne-time hardware investment
Customization capabilityLimited (can only fine-tune prompts)Fully controllable (fine-tuning, quantization)
Deployment difficultyZero deployment (use API directly)Requires technical skills to deploy
Update speedVendor continuously updatesNeed to track new versions yourself

How to Choose

Choose closed-source when:

  • You want the strongest capabilities and don't want to deal with technical details
  • Your data is not sensitive and can be sent to third parties
  • Your usage is small, pay-per-token is more cost-effective
  • You need to launch quickly and don't want to spend time on deployment

Choose open-source when:

  • Your data is highly sensitive and cannot leave the company
  • Your usage is large, pay-per-token is too expensive
  • You need deep customization of model behavior
  • You have a technical team to deploy and maintain

Recommended strategy: use closed-source for small usage and exploration (low cost, fast launch), consider open-source for large-scale and production environments (strong controllability, low long-term cost).


Cloud AI vs Local AI

This is another important choice: use cloud AI services, or run models on your own computer?

Cloud AI: Ready to Use

Cloud AI means calling AI services in the cloud via web pages or APIs.

You don't need to worry about where the model is deployed, how many GPUs are used, or how it's updated—you just input content and get results.

Typical cloud AI products:

ProductTypePayment method
ChatGPTWeb chatSubscription $20/month, or pay-as-you-go
Claude.aiWeb chatSubscription $20/month, or pay-as-you-go
OpenAI APIAPI interfaceBilled by token
Anthropic APIAPI interfaceBilled by token

Local AI: Fully Controllable

Local AI runs open-source models on your own computer or server.

Data doesn't leave your device, and there's no ongoing payment.

But you need a good graphics card—usually an NVIDIA GPU; the more VRAM the better.

Common tools for local AI:

ToolFeaturesSuitable for
OllamaOne-click start, command-line operationDevelopers, technical users
LM StudioGUI, rich model libraryGeneral users, beginners
vLLMHigh-performance inference serviceProduction environment deployment
Text Generation WebUIMost feature-complete, fine-tunableAdvanced users

Cloud AI vs Local AI Comparison

DimensionCloud AILocal AI
Initial costZero, pay as you goHigh (need to buy GPU)
Usage costContinuous pay-per-useElectricity cost + hardware depreciation
Data privacyData sent to service providerData stays entirely local
Deployment difficultyNoneRequires some technical skills
Response speedDepends on networkLocal computing, stable speed
Capability ceilingCan use the largest modelsLimited by GPU VRAM

Hybrid Strategy: Getting the Best of Both Worlds

Many companies adopt a hybrid strategy:

  • Sensitive data uses local models
  • Ordinary queries use cloud APIs
  • Use cloud services during daytime peaks, local for nighttime batch processing

This ensures data security while controlling costs.

Advice for individual users: try cloud AI first, and after getting familiar, consider local deployment if you have privacy needs or usage increases.


AI Infrastructure

AI is not just models—the underlying computing power, chips, and cloud services form the infrastructure of the entire industry.

GPU Computing Power: Why NVIDIA Dominates the Market

Training large models requires massive computation, and GPUs (Graphics Processing Units) are currently the most suitable hardware.

NVIDIA holds a near-monopoly in this field, for three reasons:

AdvantageDescription
CUDA ecosystemCUDA is NVIDIA's programming framework; almost all AI frameworks prioritize support for it.
Memory capacityProfessional cards like H100 and A100 have 80GB or even more memory.
Software supportOptimized libraries, compilers, and toolchains are the most complete.

Common NVIDIA GPU positioning:

Card modelMemoryPositioningApplicable scenarios
RTX 3090/409024GBConsumer-levelPersonal learning, small model inference
A10G24GBEntry-level professionalLightweight inference, fine-tuning
A10080GBProfessional-levelModel training, large-scale inference
H10080GBFlagship-levelLarge model training (70B+)

Memory size is a key metric.Generally speaking:

  • A 7B model requires 8-16GB of memory
  • A 13B model requires 16-32GB of memory
  • A 70B model requires 60-80GB of memory

Cloud Computing Providers: Compute as a Service

Not everyone can afford a GPU costing hundreds of thousands—cloud vendors offer a "pay-per-hour rental" model.

Cloud vendorFeaturesTypical instances
AWSMost complete instance typesp4d, p5, g5 series
AzureComplete enterprise servicesNC, ND, NV series
GCPGood TPU supportA3, T4 series
Alibaba CloudFast access in Chinagn6, gn7 series

Domestic Computing Power Alternatives

Due to chip export restrictions, China is actively developing domestic computing power.

Main players:

VendorChipFeatures
Huawei AscendAscend 910Ecosystem relatively complete, supports large model training
CambriconSiyuan 590Focused on AI acceleration
HygonDCUCUDA-ecosystem compatible

Domestic chips are still catching up; the main issue is the software ecosystem—many AI frameworks require special adaptation.

For individual developers, it's recommended to start with NVIDIA consumer-grade GPUs (e.g., RTX 4090), as the ecosystem is the most mature and has the most resources.


AI Business Models

How do AI companies make money? Currently there are three main models.

B2C: Subscription Model

Targets individual users, paid monthly.

ProductPricingWhat's included
ChatGPT Plus$20/monthPriority access to GPT-4, faster responses
Claude Pro$20/monthPriority access to Claude 3, longer context
GitHub Copilot$10/monthCode completion, code explanation

The benefit of the subscription model is stable cash flow and high user loyalty. However, it must continuously deliver value, otherwise users will cancel.

B2B: API Calls

Pay-as-you-go—you pay based on how many times you call the API.

The pricing unit is "Token," which can be understood as "word" or "character."

Typical pricing (mid-2025):

ModelInput priceOutput price
GPT-4o$5 per million tokens$15 per million tokens
Claude 3.5 Sonnet$3 per million tokens$15 per million tokens
GPT-4$30 per million tokens$60 per million tokens
Llama 3.1 405B (self-hosted)Hardware cost + electricity

Let's do the math: Suppose you have an app that processes 100,000 queries per day, with each query using 1,000 tokens.

With GPT-4o, the daily cost is $5 (input) + $15 (output) = $20, which comes to $600 per month.

If usage grows, costs rise quickly — which is why many companies ultimately choose to self-host open-source models.

Monetizing the Open Source Ecosystem

Open-source models are free, but companies can make money through peripheral services:

  • Hosting services: deploy and maintain for you
  • Enterprise edition: provide extra features and support
  • Consulting services: customize and optimize for you
  • Training services: fine-tune models with your data

Mistral AI is a representative of this model — the models are open source, but enterprises need to pay to access larger models and better services.

Comparison of the Three Models

ModelTarget usersRevenue stabilityTechnical requirements
SubscriptionIndividual usersMediumRequires excellent product experience
API callsDevelopers, enterprisesHigh (usage-driven)Requires stable service
Open-source ecosystemEnterprises, developersSlower (requires building an ecosystem)Requires community operations capability

AI Regulation and Policy

AI is developing so fast that every country is formulating regulatory policies. Understanding these policies can help you mitigate risks.

EU AI Act

The EU has the strictest regulation, with the core concept being "risk-based classification."

Risk levelDefinitionRegulatory requirements
Unacceptable riskSocial scoring, real-time facial recognitionProhibited
High riskMedical devices, legal judgments, recruitmentStrict compliance, high transparency
Medium riskChatbots, content generationLabel AI-generated content
Low riskGaming, photo editingEssentially no requirements

The EU AI Act has a global impact — as long as your product serves EU users, you must comply.

China's AI Regulatory Framework

China's regulatory policies are also being gradually refined, with core documents including:

  • "Interim Measures for the Management of Generative AI Services"
  • "Artificial Intelligence Law" (Draft for Comments)

Key requirements:

  • Generated content must be truthful and accurate, and must not fabricate false information
  • Content review mechanisms must be sound
  • Data sources must be legitimate
  • Algorithms must be transparent and explainable

US AI Policy Dynamics

US regulation is relatively lenient, primarily through executive orders and industry self-regulation.

Key focus areas:

  • National security: prevent AI from being used for malicious purposes
  • Fairness: prevent algorithmic discrimination
  • Transparency: AI-generated content needs to be labeled

For individual developers, the most important things to note are: AI-generated content should be clearly labeled and should not masquerade as human creation. When personal data is involved, data protection regulations must be followed.


2025 AI Industry Trend Outlook

As of mid-2026, several trends have become quite clear.

Trend 1: Model Capabilities Continue to Improve, but with Diminishing Marginal Returns

GPT-5, Claude 5, Gemini 3.0, DeepSeek 4, and GLM 5 have all arrived, but each upgrade brings diminishing surprises.

It's not that technology has hit a bottleneck, but rather that there are already many "good enough" use cases.

Trend 2: Open Source Models Continue to Catch Up, Narrowing the Gap with Closed Source

Llama 3.1 405B's performance is already close to GPT-4's level, and it's freely available.

More companies will join the open-source camp in the future, and the open-source model ecosystem will become increasingly mature.

Trend 3: The Application Layer Begins to Explode, Agents Become Practical

The past few years were about competition at the model layer; now the focus is shifting to the application layer.

AI Agents — AI that can autonomously complete complex tasks — will move from concept to practical application.

Trend 4: Multimodality Becomes Standard

Pure text models will become increasingly rare, and multimodal models that can handle text, images, audio, and video simultaneously will become the standard.

Trend 5: Regulation Becomes Normalized

Regulatory frameworks in various countries will gradually take shape, moving the AI industry from "reckless growth" into a stage of "regulated development."

TrendsImpact on individualsRecommendations
Model capability improvementTools are becoming more powerfulWatch for new features, but don't blindly chase novelty
Open-source models catching upLocal deployment is becoming more feasibleLearn the open-source tool stack
Explosion at the application layerOpportunities lie at the application layerCombine AI with your domain
Multimodal as standardCan do more thingsExpand multimodal application ideas
Regulatory normalizationCompliance requirements are increasingUnderstand relevant regulations

Code Example: Calling an Open Source Model API

Let's write a simple example in Python to demonstrate how to call an open-source model API.

We'll use Ollama as the local inference service, which supports multiple open-source models such as Llama 3, Qwen, and Mistral.

Reference for Ollama-related content:https://www.example.com/ollama/ollama-tutorial.html。

Example

# ============================================
# Example of calling an open-source model API
# Demo: Running the Llama 3 model locally with Ollama
# ============================================

import requests
import json


def call_ollama_model(prompt: str, model: str = "llama3") -> str:
    """
Call Ollama's API to get a model reply

Parameter descriptions:
prompt: Your question or input text
model: The model name to use (default: llama3)
Optional values: llama3, qwen2, mistral, yi, etc.

Return value:
The model's reply text
    """

    # Ollama runs on local port 11434 by default
    url = "http://localhost:11434/api/generate"

    # Build the request data
    payload = {
        "model": model,           # Model to use
        "prompt": prompt,         # User input
        "stream": False,          # Do not use streaming output (return at once)
        "options": {
            "temperature": 0.7,   # Temperature parameter: higher is more random, lower is more deterministic
            "top_p": 0.9,         # Sampling parameter: controls diversity
            "num_predict": 512,   # Maximum number of generated tokens
        }
    }

    try:
        # Send POST request
        response = requests.post(url, json=payload, timeout=60)
        response.raise_for_status()  # Check whether the request succeeded

        # Parse the returned result
        result = response.json()
        return result.get("response", "No reply obtained")

    except requests.exceptions.ConnectionError:
        return "Error: Unable to connect to the Ollama service. Please start Ollama first."
    except requests.exceptions.Timeout:
        return "Error: Request timed out. The model may still be loading."
    except Exception as e:
        return f"Error: {str(e)}"


def chat_with_example():
    """Example of multi-turn conversation with the example AI assistant"""
    print("=" * 50)
    print("🚀 example AI assistant (based on Llama 3 open-source model)")
    print("=" * 50)
    print("Tip: Enter 'quit' or 'exit' to exit the conversation")
    print()

    # Conversation history (for context continuity)
    conversation_history = []

    while True:
        # Get user input
        user_input = input("You: ").strip()

        # Check whether to exit
        if user_input.lower() in ["quit", "exit", "Exiting"]:
            print("example AI: Goodbye! Welcome back anytime.")
            break

        if not user_input:
            continue

        # Build a prompt with context
        full_prompt = "\n".join(conversation_history) + f"\nUser: {user_input}\nAssistant:"

        # Call the model
        print("example AI: ", end="", flush=True)
        response = call_ollama_model(full_prompt)
        print(response)
        print()

        # Update conversation history
        conversation_history.append(f"User: {user_input}")
        conversation_history.append(f"Assistant: {response}")

        # Keep only the last 5 turns of conversation to avoid overly long context
        if len(conversation_history) > 10:
            conversation_history = conversation_history[-10:]


# Example 1: single call
print("--- Example 1: single question ---")
question = "Please explain what a large language model is in one sentence."
print(f"Question: {question}")
answer = call_ollama_model(question)
print(f"Answer: {answer}")
print()

# Example 2: using different models
print("--- Example 2: switching between different models ---")
for model_name in ["llama3", "qwen2", "mistral"]:
    print(f"Using model {model_name}:")
    answer = call_ollama_model("Hello, please introduce yourself.", model=model_name)
    print(f"Answer: {answer[:100]}...")  # Only display the first 100 characters
print()

# Example 3: multi-turn conversation (uncomment to run)
# print("--- Example 3: multi-turn conversation ---")
# chat_with_example()

Before running this code, you need to install Ollama first:

  • gohttps://ollama.comDownload and install Ollama
  • Run in terminal:ollama pull llama3Download the model
  • Run:ollama serveStart the service (if not already started)
  • Then run the Python code above

If you don't have a local GPU, you can also use cloud-based open-source model APIs, such as Together.ai, Anyscale, etc. The calling method is similar, except the URL and authentication method are different.

The advantage of open-source models is that you can first use a small model (7B) locally for development and testing, and if it works well, switch to a larger model (70B) or cloud service.

Other extensions