text-normalization-and-large-file-processing
SkillFiles & storageCleans up Excel files by normalizing text like removing odd prefixes and extracting Chinese characters, then gives you a download link.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the text-normalization-and-large-file-processing skill
About this capability
Performs text normalization and cleaning on Excel files (e.g., removing abnormal prefixes, extracting pure Chinese characters, etc.), and finally outputs the cleaned Excel file with a download link.
What this skill tells your AI
The instructions your AI receives, as published by opensensenova/sensenova-skills in skills/sn-da-excel-workflow/capability/excel-data-cleaning/text-normalization/SKILL.md and read by ahel’s review.
Skill Steps
This sub-skill covers one capability of the Excel workflow. For reading/counting/Parquet optimization, see the parent workflow SKILL.md.
Step1 识别并清洗包含前缀符号的异常数值字段,统一转换为整数类型;同时使用正则表达式清洗文本字段,仅保留 Unicode 范围内的中文字符。
import re
import numpy as np
target_numeric_col = '需要转数字的文本列' # 示例:'获赞'
target_text_col = '需要提取中文的列' # 示例:'收货人'
# 1. 清洗包含前缀符号的数值字段
prefix_patterns = ['.', 'I ', '■ ', '一 ', '_', '. ']
def clean_numeric_with_prefix(value):
val_str = str(value).strip()
if val_str in ['None', 'nan', '', 'nan']:
return np.nan
for prefix in prefix_patterns:
if val_str.startswith(prefix):
val_str = val_str[len(prefix):].strip()
break
if val_str == '':
return np.nan
try:
return int(val_str)
except ValueError:
return np.nan
# 2. 清洗文本字段,仅保留 Unicode 范围内的中文字符(\u4e00-\u9fff)
def clean_chinese_name(name):
if pd.isna(name):
return name
s = str(name)
chinese_chars = re.findall(r'[\u4e00-\u9fff]', s)
cleaned = ''.join(chinese_chars)
return cleaned if cleaned else ''
if target_numeric_col in df.columns:
df[f'{target_numeric_col}_清洗后'] = df[target_numeric_col].apply(clean_numeric_with_prefix)
if target_text_col in df.columns:
df[f'{target_text_col}_清洗后'] = df[target_text_col].apply(clean_chinese_name)
Step2 将清洗后的结果保存为 Excel 文件,在报告中提供下载链接,并执行内存清理以应对大文件处理时的内存压力。
output_path = '/mnt/data/标准化清洗结果.xlsx'
# 保存清洗结果
df.to_excel(output_path, index=False, engine='openpyxl')
print(f'清洗结果已保存到: {output_path}')
# 生成可下载链接
print(f'[下载清洗结果表](sandbox:{output_path})')
# 内存清理
if 'df' in locals():
del df
gc.collect()
Signals
- GitHub stars
- 6k
- Forks
- 390
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
text-normalization-and-large-file-processing- Source
- github.com/opensensenova/sensenova-skills