从选题到终稿:拥抱变革,坚守学术伦理,打造高质量毕业论文
| 场景 | ❌ 错误做法 | ✅ 正确做法 |
|---|---|---|
| 选题阶段 | 让AI直接生成研究题目并直接使用 | 用AI拓展思路,结合个人兴趣确定方向 |
| 文献阅读 | 让AI总结后直接引用未读原文 | 用AI辅助筛选,精读后自己总结 |
| 数据分析 | 让AI伪造数据支持假设 | 用AI编写分析代码,基于真实数据 |
| 论文写作 | 让AI生成全文后署名提交 | 用AI辅助润色,核心内容原创 |
"数字普惠金融"
"数字普惠金融与农村消费"
"移动支付普及对农村家庭消费结构的影响研究"
import pandas as pd
import matplotlib.pyplot as plt
from wordcloud import WordCloud
import jieba
import seaborn as sns
# 设置中文字体
plt.rcParams['font.sans-serif'] = ['SimHei']
plt.rcParams['axes.unicode_minus'] = False
# 模拟文献关键词数据
keywords_data = {
'关键词': ['数字普惠金融', '农村金融', '小微企业', '收入不平等',
'金融科技', '移动支付', '绿色金融', '消费升级',
'信贷可得性', '金融排斥', '区块链', 'ESG'],
'频次': [156, 132, 128, 115, 108, 95, 88, 82, 76, 71, 65, 58],
'年份': [2023] * 12
}
df_keywords = pd.DataFrame(keywords_data)
# 绘制高频关键词条形图
fig, axes = plt.subplots(1, 2, figsize=(14, 6))
# 子图1:关键词频次柱状图
ax1 = axes[0]
top_keywords = df_keywords.iloc[:12]
colors = sns.color_palette("husl", 12)
bars = ax1.barh(top_keywords['关键词'], top_keywords['频次'], color=colors)
ax1.set_xlabel('频次', fontsize=12)
ax1.set_title('普惠金融研究高频关键词 (2023)', fontsize=14, fontweight='bold')
ax1.invert_yaxis()
# 添加数值标签
for bar in bars:
width = bar.get_width()
ax1.text(width + 3, bar.get_y() + bar.get_height()/2,
f'{int(width)}', ha='left', va='center', fontsize=10)
# 子图2:词云图
ax2 = axes[1]
word_freq = dict(zip(df_keywords['关键词'], df_keywords['频次']))
wordcloud = WordCloud(width=500, height=400,
background_color='white',
font_path='SimHei.ttf',
colormap='viridis').generate_from_frequencies(word_freq)
ax2.imshow(wordcloud, interpolation='bilinear')
ax2.axis('off')
ax2.set_title('关键词词云图', fontsize=14, fontweight='bold', pad=20)
plt.tight_layout()
plt.savefig('literature_keywords.png', dpi=300, bbox_inches='tight')
plt.show()
print("✅ 文献计量分析完成!")
print(f"📊 热度最高的关键词:{df_keywords.iloc[0]['关键词']}")
方法1:知网/万方
在文献详情页点击Zotero插件图标,自动抓取元数据
方法2:手动添加
手动输入ISBN或DOI,Zotero自动获取文献信息
方法1:Google Scholar
设置中添加Zotero,直接点击"Import into Zotero"
方法2:Web of Science
批量选中文献后导出为RIS格式导入Zotero
Word文档:安装Zotero Word插件,光标位置点击"Add/Edit Citation"插入引用
LaTeX文档:右键文献库→导出→选择BibTeX格式→保存为references.bib
| 场景 | 正确格式 | 说明 |
|---|---|---|
| 参考文献列表 | Smith, J. A. | 姓在前,名在后用缩写 |
| 正文引用 | Smith (2023) 或 (Smith, 2023) | 只用姓 |
| 两位作者 | Smith, J. A., & Doe, B. B. | &前有逗号 |
| 三位及以上 | Smith, J. A., et al. | 第一作者+et al. |
| 中文作者拼音 | Zhang, S. | 姓全大写,名首字母大写 |
import pandas as pd
import numpy as np
from scipy import stats
# ============ 数据加载 ============
print("📂 加载数据...")
df = pd.read_csv('digital_finance_panel.csv')
print(f"原始数据形状: {df.shape}")
# ============ 缺失值处理 ============
print("\n🔍 检查缺失值...")
missing_stats = df.isnull().sum()
missing_pct = (missing_stats / len(df) * 100).round(2)
print("缺失值统计:")
for col in missing_pct[missing_pct > 0].index:
print(f" {col}: {missing_pct[col]}%")
# 删除缺失率超过20%的变量
df = df.dropna(thresh=len(df)*0.8, axis=1)
# 数值变量用中位数填补
num_cols = df.select_dtypes(include=[np.number]).columns
df[num_cols] = df[num_cols].fillna(df[num_cols].median())
# ============ 异常值处理 ============
print("\n⚠️ 处理异常值...")
def remove_outliers(df, column, method='iqr'):
if method == 'iqr':
Q1 = df[column].quantile(0.25)
Q3 = df[column].quantile(0.75)
IQR = Q3 - Q1
lower = Q1 - 1.5 * IQR
upper = Q3 + 1.5 * IQR
df_clean = df[(df[column] >= lower) & (df[column] <= upper)]
removed = len(df) - len(df_clean)
print(f" {column}: 移除 {removed} 个异常值")
return df_clean
# 对关键变量进行异常值处理
key_vars = ['digital_index', 'household_income', 'consumption']
for var in key_vars:
df = remove_outliers(df, var, method='iqr')
# ============ 描述性统计 ============
print("\n📊 描述性统计:")
print(df.describe().round(4))
df.to_csv('data_cleaned.csv', index=False)
print(f"\n✅ 数据清洗完成!最终数据形状: {df.shape}")
| digital_index | household_income | consumption | |
| count | 12237.0000 | 12237.0000 | 12237.0000 |
| mean | 185.2341 | 45678.1234 | 32456.7891 |
| std | 52.3412 | 12345.6789 | 9876.5432 |
| min | 89.1234 | 12345.0000 | 9876.0000 |
import statsmodels.formula.api as smf
from linearmodels.panel import PanelOLS
# ============ 基准回归:固定效应模型 ============
print("🔬 基准回归分析...")
# 个体固定效应模型
model_fe = smf.ols(
'household_income ~ digital_index + gdp_per_capita + '
'education + C(province) + C(year)',
data=df
).fit(cov_type='cluster', cov_kwds={'groups': df['province']})
print("\n📊 固定效应模型结果:")
print(model_fe.summary())
# ============ 稳健性检验 ============
print("\n🔍 稳健性检验...")
# 检验1: 滞后解释变量
df['digital_index_lag'] = df.groupby('province')['digital_index'].shift(1)
model_robust1 = smf.ols(
'household_income ~ digital_index_lag + gdp_per_capita + '
'education + C(province) + C(year)',
data=df.dropna()
).fit()
# 检验2: 剔除直辖市
df_no_muni = df[~df['province'].isin(['北京', '上海', '天津', '重庆'])]
model_robust2 = smf.ols(
'household_income ~ digital_index + gdp_per_capita + '
'education + C(province) + C(year)',
data=df_no_muni
).fit()
# 检验3: 安慰剂检验
np.random.seed(42)
df['digital_placebo'] = np.random.permutation(df['digital_index'])
model_placebo = smf.ols(
'household_income ~ digital_placebo + gdp_per_capita + '
'education + C(province) + C(year)',
data=df
).fit()
print("\n✅ 稳健性检验完成!")
print(f"基准模型系数: {model_fe.params['digital_index']:.4f}***")
print(f"稳健性检验1系数: {model_robust1.params['digital_index_lag']:.4f}***")
print(f"稳健性检验2系数: {model_robust2.params['digital_index']:.4f}***")
print(f"安慰剂检验系数: {model_placebo.params['digital_placebo']:.4f} (应不显著)")
import matplotlib.pyplot as plt
import seaborn as sns
import pandas as pd
import numpy as np
# 设置学术风格
plt.style.use('seaborn-v0_8-whitegrid')
sns.set_palette("husl")
# 设置中文字体
plt.rcParams['font.sans-serif'] = ['SimHei']
plt.rcParams['axes.unicode_minus'] = False
# ============ 图1: 数字普惠金融指数时空演变热力图 ============
fig, axes = plt.subplots(1, 2, figsize=(14, 5))
# 子图1: 各省份数字金融指数热力图
provinces = ['北京', '上海', '广东', '浙江', '江苏', '四川']
years = [2015, 2018, 2021, 2023]
data_heatmap = np.array([
[245, 312, 398, 456],
[238, 298, 385, 442],
[198, 256, 334, 398],
[205, 268, 345, 412],
[192, 248, 328, 389],
[165, 218, 298, 356]
])
im = axes[0].imshow(data_heatmap, cmap='RdYlGn', aspect='auto')
axes[0].set_xticks(range(len(years)))
axes[0].set_xticklabels(years)
axes[0].set_yticks(range(len(provinces)))
axes[0].set_yticklabels(provinces)
axes[0].set_title('各省份数字普惠金融指数演变', fontweight='bold')
# 添加数值标签
for i in range(len(provinces)):
for j in range(len(years)):
text = axes[0].text(j, i, f'{data_heatmap[i,j]}',
ha="center", va="center", fontsize=9)
plt.colorbar(im, ax=axes[0], label='指数值')
# 子图2: 东中西部地区差异趋势
axes[1].plot(years, data_heatmap[0], marker='o', label='北京')
axes[1].plot(years, data_heatmap[4], marker='s', label='江苏')
axes[1].plot(years, data_heatmap[5], marker='^', label='四川')
axes[1].set_title('区域数字金融发展对比', fontweight='bold')
axes[1].set_xlabel('年份')
axes[1].set_ylabel('数字普惠金融指数')
axes[1].legend()
axes[1].grid(True, alpha=0.3)
plt.tight_layout()
plt.savefig('figure1_spatial_temporal.png', dpi=300)
print("✅ 图1保存成功!")
# ============ 图2: 回归系数森林图 ============
fig, ax = plt.subplots(figsize=(10, 6))
models = ['基准模型', '稳健性检验1', '稳健性检验2', '稳健性检验3']
coefficients = [0.152, 0.135, 0.168, 0.141]
ci_lower = [0.110, 0.095, 0.128, 0.105]
ci_upper = [0.194, 0.175, 0.208, 0.177]
y_pos = np.arange(len(models))
ax.errorbar(coefficients, y_pos,
xerr=[[(c-l) for c,l in zip(coefficients, ci_lower)],
[(u-c) for c,u in zip(coefficients, ci_upper)]],
fmt='o', markersize=10, capsize=5, linewidth=2.5,
color='#007bff', ecolor='#6f42c1')
ax.set_yticks(y_pos)
ax.set_yticklabels(models)
ax.set_xlabel('系数值')
ax.set_title('数字普惠金融对家庭收入的影响:稳健性检验', fontweight='bold')
ax.axvline(x=0, color='black', linestyle='--', linewidth=1.5)
ax.grid(True, alpha=0.3, axis='x')
plt.tight_layout()
plt.savefig('figure2_forest_plot.png', dpi=300)
print("✅ 图2保存成功!")
论点句
清晰表达核心观点
理论支撑
引用相关理论/文献
证据陈述
展示数据/经验证据
解释说明
阐述证据与论点关系
意义总结
点明研究价值
{
"latex-workshop.latex.autoBuild.run": "onSave",
"latex-workshop.view.pdf.viewer": "tab",
"latex-workshop.latex.tools": [
{
"name": "xelatex",
"command": "xelatex",
"args": ["-synctex=1", "-interaction=nonstopmode", "%DOC%"]
}
]
}
\documentclass[12pt,a4paper]{article}
\usepackage[UTF8]{ctex} % 中文支持
\usepackage{amsmath,amssymb} % 数学公式
\usepackage{booktabs} % 三线表
\usepackage{graphicx} % 图片
\usepackage{cite} % 参考文献
\usepackage{hyperref} % 超链接
\title{数字普惠金融对农村家庭消费的影响研究}
\author{张三 \\ 某某大学金融学院}
\date{\today}
\begin{document}
\maketitle
\section{引言}
数字普惠金融作为金融创新的重要形式,对促进农村消费升级具有重要意义\cite{zhang2023}。
\section{模型设定}
基准回归模型设定如下:
\begin{equation}
\ln(Consumption_{it}) = \alpha_0 + \alpha_1 Digital_{it} + \alpha_2 X_{it} + \mu_i + \lambda_t + \varepsilon_{it}
\label{eq:baseline}
\end{equation}
其中,下标$i$和$t$分别表示省份和年份;$Digital_{it}$为数字普惠金融指数;$X_{it}$为控制变量向量;$\mu_i$和$\lambda_t$分别为省份和年份固定效应。
\section{实证结果}
表\ref{tab:baseline}报告了基准回归结果。
\begin{table}[htbp]
\centering
\caption{基准回归结果}
\label{tab:baseline}
\begin{tabular}{lcccc}
\toprule
& \multicolumn{2}{c}{固定效应模型} & \multicolumn{2}{c}{随机效应模型} \\
\cmidrule(lr){2-3} \cmidrule(lr){4-5}
变量 & 系数 & 标准误 & 系数 & 标准误 \\
\midrule
数字普惠金融指数 & 0.152*** & (0.021) & 0.148*** & (0.019) \\
人均GDP(对数) & 0.089** & (0.035) & 0.092** & (0.033) \\
受教育年限 & 0.045*** & (0.012) & 0.043*** & (0.011) \\
\midrule
省份固定效应 & Yes & & No & \\
年份固定效应 & Yes & & Yes & \\
观测数 & 12,450 & & 12,450 & \\
$R^2$ & 0.312 & & 0.298 & \\
\bottomrule
\end{tabular}
\medskip
\\ \footnotesize 注:***、**、*分别表示在1%、5%、10%水平上显著。
\end{table}
从表\ref{tab:baseline}可以看出,数字普惠金融指数的系数为0.152,在1\%水平上显著,表明数字普惠金融每提高1个单位,农村家庭消费支出增加约15.2\%。
\bibliographystyle{plain}
\bibliography{references}
\end{document}
张三
某某大学金融学院
2024年3月5日
数字普惠金融作为金融创新的重要形式,对促进农村消费升级具有重要意义[1]。
基准回归模型设定如下:
其中,下标i和t分别表示省份和年份;Digitalit为数字普惠金融指数;Xit为控制变量向量;μi和λt分别为省份和年份固定效应。
表1报告了基准回归结果。
| 固定效应模型 | 随机效应模型 | |||
|---|---|---|---|---|
| 变量 | 系数 | 标准误 | 系数 | 标准误 |
| 数字普惠金融指数 | 0.152*** | (0.021) | 0.148*** | (0.019) |
| 人均GDP(对数) | 0.089** | (0.035) | 0.092** | (0.033) |
| 受教育年限 | 0.045*** | (0.012) | 0.043*** | (0.011) |
| R² | 0.312 | 0.298 | ||
注:***、**、*分别表示在1%、5%、10%水平上显著。
从表1可以看出,数字普惠金融指数的系数为0.152,在1%水平上显著,表明数字普惠金融每提高1个单位,农村家庭消费支出增加约15.2%。
# ============ 初始化仓库 ============
git init # 初始化本地仓库
git config user.name "你的名字"
git config user.email "你的邮箱"
# ============ 连接远程仓库 ============
git remote add origin https://gitee.com/用户名/thesis-2024.git
# 或GitHub: git remote add origin https://github.com/用户名/thesis-2024.git
# ============ 基本工作流程 ============
git add . # 添加所有修改
git status # 查看状态
git commit -m "完成第二章初稿" # 提交修改(附说明)
git push origin main # 推送到远程仓库
# ============ 常用操作 ============
git log --oneline # 查看提交历史
git diff # 查看未暂存的修改
git checkout -- 文件名.tex # 撤销文件修改
git reset --hard HEAD~1 # 回退到上一个版本
# ============ .gitignore 配置 ============
# 创建.gitignore文件,忽略以下文件:
*.aux # LaTeX辅助文件
*.log # 日志文件
*.out # 输出文件
*.toc # 目录文件
*.pdf # 可选:忽略生成的PDF
~$* # Word临时文件
.DS_Store # Mac系统文件
| 特性 | Notion | Obsidian |
|---|---|---|
| 数据存储 | 云端存储,依赖网络 | 本地Markdown文件,完全掌控 |
| 学习曲线 | 较低,界面直观 | 较高,需熟悉插件系统 |
| 数据库功能 | ✅ 强大的数据库视图 | ❌ 需插件实现 |
| 可扩展性 | 有限,依赖官方功能 | ✅ 丰富的插件生态 |
| 离线使用 | ❌ 需网络同步 | ✅ 完全离线可用 |
| 版本控制 | 内置页面历史 | ✅ 兼容Git版本控制 |
| 适合场景 | 团队协作、项目管理 | 个人知识库、学术笔记 |
💡 提示:可以将以下代码复制到 mermaid.live 在线编辑器中实时预览和编辑!
```mermaid
graph TD
A["数据来源"] --> A1["CFPS 中国家庭追踪调查"]
A --> A2["CHFS 中国家庭金融调查"]
A --> A3["北大数字普惠金融指数"]
A --> A4["统计年鉴宏观数据"]
A1 --> B["数据清洗与处理"]
A2 --> B
A3 --> B
A4 --> B
B --> B1["缺失值处理(删除/插值)"]
B --> B2["异常值处理(IQR方法)"]
B --> B3["数据合并(面板数据构建)"]
B1 --> C["变量选择"]
B2 --> C
B3 --> C
C --> C1["核心变量:数字普惠金融指数"]
C --> C2["被解释变量:家庭收入/消费"]
C --> C3["控制变量:GDP/教育/年龄"]
C --> C4["机制变量:创业/信贷可得性"]
C1 --> D["模型构建"]
C2 --> D
C3 --> D
C4 --> D
D --> D1["基准回归(OLS/固定效应)"]
D --> D2["工具变量法(IV/2SLS)"]
D --> D3["双重差分(DID)"]
D --> D4["中介效应模型"]
D1 --> E["实证分析"]
D2 --> E
D3 --> E
D4 --> E
E --> E1["描述性统计(均值/标准差/相关系数)"]
E --> E2["基准回归(系数显著性)"]
E --> E3["平行趋势检验(DID前)"]
E1 --> F{"稳健性检验"}
E2 --> F
E3 --> F
F --> F1["替换变量(滞后一期)"]
F --> F2["改变样本(剔除直辖市)"]
F --> F3["PSM 倾向得分匹配"]
F --> F4["安慰剂检验(随机化处理)"]
F1 --> G["异质性分析"]
F2 --> G
F3 --> G
F4 --> G
G --> G1["城乡差异"]
G --> G2["区域差异(东中西部)"]
G --> G3["收入分组(低中高)"]
G --> G4["教育水平"]
G1 --> H["机制检验"]
G2 --> H
G3 --> H
G4 --> H
H --> H1["中介效应(Bootstrap检验)"]
H --> H2["调节效应(交互项分析)"]
H --> H3["渠道分析(分步回归)"]
H1 --> I["研究结论与政策建议"]
H2 --> I
H3 --> I
I --> I1["核心结论"]
I --> I2["政策建议"]
I --> I3["研究局限"]
I --> I4["未来研究方向"]
classDef dataSource fill:#e3f2fd,stroke:#1976d2,stroke-width:2px
classDef dataProcess fill:#bbdefb,stroke:#1565c0,stroke-width:2px
classDef variable fill:#f3e5f5,stroke:#7b1fa2,stroke-width:2px
classDef model fill:#e8f5e9,stroke:#388e3c,stroke-width:2px
classDef analysis fill:#fff3e0,stroke:#f57c00,stroke-width:2px
classDef robust fill:#fce4ec,stroke:#c2185b,stroke-width:2px
classDef hetero fill:#e0f2f1,stroke:#00796b,stroke-width:2px
classDef mechanism fill:#f1f8e9,stroke:#558b2f,stroke-width:2px
classDef conclusion fill:#e8eaf6,stroke:#4527a0,stroke-width:2px
class A,A1,A2,A3,A4 dataSource
class B,B1,B2,B3 dataProcess
class C,C1,C2,C3,C4 variable
class D,D1,D2,D3,D4 model
class E,E1,E2,E3 analysis
class F,F1,F2,F3,F4 robust
class G,G1,G2,G3,G4 hetero
class H,H1,H2,H3 mechanism
class I,I1,I2,I3,I4 conclusion
```
在Excalidraw中,点击右上角导出按钮,可选择:
建议导出为SVG格式,可在LaTeX中高质量使用,或导出为PNG直接插入Word。
# API端点配置
API Base URL: https://api.deepseek.com
API Key: sk-xxxxxxxxxxxxxxxxxxxxxxxxx
Model: deepseek-chat
# 在Copilot插件中的配置
Provider Name: Deepseek
API Endpoint: https://api.deepseek.com/v1
API Key: [你的API Key]
Model: deepseek-chat
import pandas as pd
import stata_setup
# ============ 配置Stata路径 ============
# 首次使用需要配置Stata安装路径
stata_setup.config("C:/Program Files/Stata17", "se")
# 导入Stata API
from pystata import stata
# ============ 加载数据到Stata ============
print("📂 加载数据到Stata...")
df = pd.read_csv('data_cleaned.csv')
# 将Pandas DataFrame导入Stata
stata.df_store(df, "mydata")
# ============ 在Stata中执行命令 ============
print("\n🔬 执行固定效应回归...")
# 设置面板数据
stata.run("tsset province_id year")
# 固定效应回归
stata.run("""
xtreg household_income digital_index gdp_per_capita education, fe
cluster(province)
""", quietly=False)
# 保存回归结果
stata.run("estimates store fe_model")
# ============ 生成三线表(esttab) ============
print("\n📊 生成LaTeX格式结果表...")
stata.run("""
esttab fe_model using "regression_results.tex", ///
replace b(3) se(3) star(* 0.1 ** 0.05 *** 0.01) ///
scalars("r2_a R-squared N N") ///
title("数字普惠金融对家庭收入的影响")
""")
# ============ 导出结果到Python ============
# 获取回归系数
coef_df = stata.get_df()
print("\n✅ 回归完成!")
print(coef_df.head())
# ============ 批量回归示例 ============
print("\n🔄 批量回归不同模型...")
models = [
"household_income digital_index",
"household_income digital_index gdp_per_capita",
"household_income digital_index gdp_per_capita education",
"household_income digital_index gdp_per_capita education i.province i.year"
]
for i, model in enumerate(models):
print(f"模型 {i+1}: {model}")
stata.run(f"reg {model}")
stata.run(f"estimates store model{i+1}")
# 比较所有模型
stata.run("esttab model* using "all_models.tex", replace star(* 0.1 ** 0.05 *** 0.01)")
print("\n✅ 批量回归完成!结果已保存。")
| 功能 | Stata命令 | Python等效操作 |
|---|---|---|
| 描述性统计 | summarize 或 sum |
df.describe() |
| 相关性矩阵 | correlate 或 cor |
df.corr() |
| OLS回归 | reg y x1 x2 |
smf.ols('y~x1+x2').fit() |
| 固定效应 | xtreg y x, fe |
PanelOLS().fit() |
| 工具变量 | ivregress 2sls y (x=z) |
IV2SLS().fit() |
| 导出结果 | esttab |
summary_col() |
| 特性 | CSV文件 | 数据库(PostgreSQL) |
|---|---|---|
| 数据规模 | ❌ 超过1GB性能下降 | ✅ 支持TB级数据 |
| 查询速度 | ❌ 全表扫描 | ✅ 索引加速查询 |
| 并发访问 | ❌ 文件锁定 | ✅ 多用户并发 |
| 数据完整性 | ❌ 无约束 | ✅ ACID事务保证 |
| 关联查询 | ❌ 需手动编写 | ✅ SQL JOIN支持 |
| 增量更新 | ❌ 需重写文件 | ✅ 按行更新 |
SQL (Structured Query Language) 是结构化查询语言,用于管理关系型数据库。
import psycopg2
import pandas as pd
from sqlalchemy import create_engine
# ============ 方法1:使用psycopg2 ============
print("📡 连接PostgreSQL数据库...")
conn = psycopg2.connect(
host="localhost",
database="finance_db",
user="postgres",
password="your_password"
)
cursor = conn.cursor()
# ============ 创建表 ============
print("\n📊 创建数据表...")
cursor.execute("""
CREATE TABLE IF NOT EXISTS digital_finance (
id SERIAL PRIMARY KEY,
province VARCHAR(50),
year INTEGER,
digital_index FLOAT,
household_income FLOAT,
gdp_per_capita FLOAT,
education FLOAT
)
""")
conn.commit()
# ============ 插入数据 ============
print("💾 插入数据...")
cursor.execute("""
INSERT INTO digital_finance
(province, year, digital_index, household_income, gdp_per_capita, education)
VALUES (%s, %s, %s, %s, %s, %s)
""", ("北京", 2023, 456.78, 89234.5, 189000.0, 12.5))
conn.commit()
# ============ 查询数据 ============
print("\n🔍 查询数据...")
cursor.execute("""
SELECT province, year, digital_index, household_income
FROM digital_finance
WHERE year >= 2020
ORDER BY digital_index DESC
LIMIT 5
""")
results = cursor.fetchall()
for row in results:
print(f"{row[0]}({row[1]}): 指数={row[2]:.2f}, 收入={row[3]:.1f}")
# ============ 方法2:使用SQLAlchemy + Pandas ============
print("\n🐼 使用SQLAlchemy读取数据...")
engine = create_engine('postgresql+psycopg2://postgres:password@localhost/finance_db')
# 读取SQL查询到DataFrame
df = pd.read_sql_query("""
SELECT
province,
AVG(digital_index) as avg_index,
AVG(household_income) as avg_income
FROM digital_finance
GROUP BY province
ORDER BY avg_index DESC
""", engine)
print("\n📊 各省份平均数据:")
print(df.head())
# ============ 聚合查询示例 ============
print("\n📈 年度趋势分析...")
yearly_stats = pd.read_sql_query("""
SELECT
year,
COUNT(*) as obs_count,
AVG(digital_index) as avg_digital,
STDDEV(digital_index) as sd_digital,
MIN(digital_index) as min_digital,
MAX(digital_index) as max_digital
FROM digital_finance
GROUP BY year
ORDER BY year
""", engine)
print(yearly_stats)
# ============ 关联查询(JOIN)示例 ============
print("\n🔗 多表关联查询...")
# 假设有另一个省份信息表
join_result = pd.read_sql_query("""
SELECT
d.province,
d.year,
d.digital_index,
d.household_income,
p.region_name,
p.gdp_level
FROM digital_finance d
LEFT JOIN province_info p
ON d.province = p.province_name
WHERE d.year = 2023
ORDER BY d.digital_index DESC
""", engine)
print(join_result.head())
# 关闭连接
cursor.close()
conn.close()
print("\n✅ 数据库操作完成!")
| province | 北京 | 上海 | 广东 |
| avg_index | 423.5 | 415.8 | 365.2 |
| avg_income | 85234 | 88956 | 65432 |
| year | obs_count | avg_digital | sd_digital |
| 2020 | 30 | 245.3 | 52.1 |
| 2021 | 30 | 298.7 | 58.4 |
| 2023 | 30 | 365.2 | 65.8 |
SELECT column1, column2 FROM table_name WHERE condition ORDER BY column1 DESC LIMIT 10;
SELECT
province,
AVG(value) as avg_val,
COUNT(*) as count,
MAX(value) as max_val
FROM table
GROUP BY province;
SELECT a.*, b.col2 FROM table_a a LEFT JOIN table_b b ON a.id = b.a_id WHERE a.year = 2023;
-- 插入数据 INSERT INTO table VALUES (...); -- 更新数据 UPDATE table SET col = val WHERE id = 1; -- 删除数据 DELETE FROM table WHERE id = 1;
AI是强大的副驾驶(Co-pilot),而非代笔者(Ghostwriter)。毕业论文的灵魂在于原创性思考。
诚实声明AI使用,绝不伪造数据或冒充AI内容为原创思想。学术诚信是不可逾越的底线。
将AI融入选题→文献→实证→写作→排版全流程,实现研究效率与质量的双重提升。
AI工具日新月异,保持学习心态,掌握最新的Prompt技巧和工具组合。
选择1-2个方向深入学习,避免贪多嚼不烂。推荐从知识管理和计量工具开始。