Add HeXi bot codebase: custom plugins, web frontends, tests

- hexi core: message handling, rate limiting, cooldown, plugin manager
- Custom plugins: BF stats, daily check-in, quotes, persona cards, etc.
- Community plugins vendored under hexi/plugins with local fixes
- Web admin frontends (learning-chat, persona-admin), unified hexi/web
- Tests for rate_limit/cooldown/memes/persona; poetry.lock

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
2026-09-01 13:13:40 +08:00
co-authored by Claude
parent 1783c60afa
commit b61d09f09f
3201 changed files with 160436 additions and 171 deletions
@@ -0,0 +1,76 @@
# picfinder_take
个人~~缝合~~修改的Hoshino bot搜图插件。
这个插件的主要思路其实是在hoshino上还原隔壁 [@Tsuk1ko](https://github.com/Tsuk1ko) 大佬家的[竹竹](https://github.com/Tsuk1ko/cq-picsearcher-bot)的搜图交互体验(所以插件名是たけ)
代码主体部分参考了 [@Watanabe-Asa](https://github.com/Watanabe-Asa)大佬的 [搜图](https://github.com/pcrbot/Salmon-plugin-transplant#%E6%90%9C%E5%9B%BE)与 [@Cappuccilo](https://github.com/Cappuccilo)大佬的 [以图搜图](https://github.com/pcrbot/cappuccilo_plugins#%E4%BB%A5%E5%9B%BE%E6%90%9C%E5%9B%BE),感谢各位大佬的代码)
---
- 6.1更新:增加私聊搜图、截屏识别和代理功能
- 6.24更新:增加回复搜图功能,搜图请求更换为异步(感谢 [@蓝红心](https://github.com/LHXnois)大佬)
- 7.9更新:增加自定义HOST功能,在qiang内又不想用全局代理的用户(?)可以单独为SauceNao和ascii2d配置反代或ip直连)(继续感谢 [@蓝红心](https://github.com/LHXnois)大佬)
- 7.26更新:回复搜图增加at支持
- 10.20更新:增加批量搜图计数与吞图提示
- 12.12更新:增加频道搜图支持,适配go-cqhttp-1.0.0beat8-fix2(还是感谢 [@蓝红心](https://github.com/LHXnois)大佬)
- 22.01.10更新:适配新版图片CQcode,增加IGNORE_STAMP配置项,可在批量搜索中自动忽略subType不为0的图片(比如表情包);用cloudscraper绕过ascii2d近期新增的cf认证
## 特点
- 搜索SauceNao,在相似率过低时自动补充搜索ascii2d,相似率阈值可在config中调整。搜索结果显示数量可在config中调整。
- 解析SauceNao和ascii2d搜索结果的作品详细信息。SauceNao全部42个index格式解析都已完成;ascii2d常见格式应该也能解析,一些奇怪格式的外部登录不敢保证)
- 获取SauceNao和ascii2d的结果缩略图,缩略图可在config中关闭。
- 搜图结果可由普通回复切换为合并转发回复,减少刷屏情况。发送模式可在config中调整。
- 增加批量搜图模式,解决移动端文字命令+图片发送麻烦的问题。
- 增加搜图每人每日限额,可在config中调整限额。要懂得节制噢.jpg
- 增加简单的手机截屏识别功能,判断为整屏手机截屏时会拒绝搜索 ~~(你会截你马个图.jpg)~~
- 增加私聊搜图功能,有效缓解腾讯吞图(但反之临时会话下搜索结果含图片与链接时概率被吞无法发送,若要稳定使用需加bot好友)
- 增加代理设置,方便qiang内使用
- 增加回复搜图功能,可直接对群友发送的图片进行回复搜索,省去转发过程)
- New!增加IGNORE_STAMP配置项,可在批量搜索中自动忽略沙雕群友发的表情包,减少资源浪费与刷屏)
## 用法
- 申请并在config中配置SauceNao的API key
- 发送 ``@bot+图片`` 或 ``[bot昵称]搜图+图片`` 进行单张搜索。
- 发送 ``[bot昵称]搜图`` 进入连续搜图模式,连续搜图模式下同一用户所发送所有图片都将直接搜索;
发送 ``谢谢[bot昵称]`` 退出连续搜图模式,或停止发图等待超时后自动退出连续搜图模式。
- 对他人发送的图片回复 ``@bot 搜图`` 或``[bot昵称]搜图`` 进行回复搜索。
- 私聊下直接发送图片即可进行搜索。
## 频道使用
对特定子频道中收到的所有图片进行搜索(本质应该算是个常驻的连续搜图?)
在qq群这样可能影响讨论,但是在频道可以专门开一个搜图子频道只用来搜图!
食用方法:
- 创建一个专门的子频道
- 搞到频道id和子频道id(看log)
- 按{频道id: [子频道id]}格式填到config.py里的enableguild里
- 发送图片开始搜图
开启慢速模式然后给bot管理体验更佳
@@ -0,0 +1,466 @@
import re
from asyncio import sleep
from datetime import datetime, timedelta
from nonebot import get_driver, on_message, on_startswith, require
from nonebot.adapters.onebot.v11 import Bot, GroupMessageEvent, PrivateMessageEvent
from nonebot.exception import ActionFailed
from nonebot.log import logger
from nonebot.plugin import PluginMetadata
from nonebot.rule import is_type, to_me
require("nonebot_plugin_alconna")
from nonebot_plugin_alconna import UniMessage
from .config import (
CHAIN_REPLY,
CHECK,
DAILY_LIMIT,
IGNORE_STAMP,
SAUCENAO_KEY,
SEARCH_TIMEOUT,
threshold,
)
from .image import check_screenshot, get_image_data_ascii, get_image_data_sauce
# 统一配置标准:把「插件内部常量」也暴露给 Web 管理台(来源无关)。
# 说明:Web 读取/修改会写入 plugin_config.json 并热更新 config 模块属性;
# 要在运行期立即生效,插件各使用点应改用 get_effective_value 读取(见 docs/plugin-config-standard.md §5b)。
from hexi.config_standard import register_config_items # noqa: E402
from . import config as _cfgmod # noqa: E402
register_config_items(
__name__,
[
{"key": "SAUCENAO_KEY", "label": "SauceNAO API Key", "type": "password", "secret": True},
{"key": "DAILY_LIMIT", "label": "每日搜图限额", "type": "int", "default": 50},
{"key": "SEARCH_TIMEOUT", "label": "批量搜索超时(秒)", "type": "int", "default": 60},
{"key": "CHECK", "label": "开启截屏判定", "type": "bool", "default": True},
{"key": "THUMB_ON", "label": "启用缩略图", "type": "bool", "default": True},
{"key": "CHAIN_REPLY", "label": "合并转发回复", "type": "bool", "default": True},
{"key": "IGNORE_STAMP", "label": "批量忽略表情包", "type": "bool", "default": True},
{"key": "threshold", "label": "相似度阈值", "type": "float", "default": 70},
],
store=_cfgmod,
)
__plugin_meta__ = PluginMetadata(
name="搜图",
description="SauceNAO + ascii2d 二次元搜图",
usage=(
"(xx为bot名称)\n"
"[@bot+图片] 单张/多张搜图\n"
"[xx搜图] 进入批量搜图模式\n"
"[谢谢xx] 退出批量搜图模式"
),
type="application",
)
driver_config = get_driver().config
SUPERUSERS = driver_config.superusers
NICKNAMES = list(driver_config.nickname) if driver_config.nickname else ["竹竹"]
NICKNAME = NICKNAMES[0]
class DailyNumberLimiter:
"""每日次数限制器(按年积日计数)"""
def __init__(self, max_num):
self.max = max_num
self.today = -1
self.count = {}
def check(self, key) -> bool:
day = datetime.now().timetuple().tm_yday
if day != self.today:
self.today = day
self.count.clear()
return self.count.get(key, 0) < self.max
def get_num(self, key):
return self.count.get(key, 0)
def increase(self, key, num=1):
self.count[key] = self.count.get(key, 0) + num
lmtd = DailyNumberLimiter(DAILY_LIMIT)
class PicListener:
def __init__(self):
self.on = {}
self.count = {}
self.limit = {}
self.timeout = {}
def get_on_off_status(self, gid):
return self.on[gid] if self.on.get(gid) is not None else False
def turn_on(self, gid, uid):
self.on[gid] = uid
self.timeout[gid] = datetime.now() + timedelta(seconds=SEARCH_TIMEOUT)
self.count[gid] = 0
self.limit[gid] = DAILY_LIMIT - lmtd.get_num(uid)
def turn_off(self, gid):
self.on.pop(gid)
self.count.pop(gid)
self.timeout.pop(gid)
self.limit.pop(gid)
def count_plus(self, gid):
self.count[gid] += 1
pls = PicListener()
start_finder = on_startswith(("识图", "搜图", "查图", "找图"), rule=to_me())
picmessage = on_message(rule=is_type(GroupMessageEvent), block=False)
replymessage = on_message(rule=is_type(GroupMessageEvent), block=False)
thanks = on_startswith("谢谢")
pic_private = on_message(rule=is_type(PrivateMessageEvent), block=False)
def parse_report(text: str) -> UniMessage:
"""将含 [CQ:image,file=...] 的搜索报告字符串转为 UniMessage"""
msg = UniMessage()
parts = re.split(r"(\[CQ:image,[^\]]*\])", text)
for part in parts:
if part.startswith("[CQ:image,"):
m = re.search(r"file=([^,\]]+)", part)
msg += UniMessage.image(url=m.group(1) if m else "")
elif part:
msg += UniMessage.text(part)
return msg
def _extract_image(ev) -> tuple[str, str, str | None] | None:
"""从消息中提取第一张图片的 (file, url, subType),无图片返回 None"""
for m in ev.message:
if m.type == "image":
return m.data["file"], m.data["url"], m.data.get("subType")
return None
@start_finder.handle()
async def start_finder_handle(bot: Bot, ev: GroupMessageEvent):
uid = ev.user_id
gid = ev.group_id
mid = ev.message_id
if str(uid) not in SUPERUSERS:
await UniMessage.text("暂不支持在群聊中使用!").send()
return
ret = _extract_image(ev)
if not ret:
if pls.get_on_off_status(gid):
if uid == pls.on[gid]:
pls.timeout[gid] = datetime.now() + timedelta(seconds=30)
await UniMessage.text(
"您已经在搜图模式下啦!\n如想退出搜图模式请发送“谢谢竹竹”~"
).send()
await start_finder.finish()
else:
await UniMessage.at(user_id=str(pls.on[gid])).text(
"正在搜图,请耐心等待~"
).send()
await start_finder.finish()
pls.turn_on(gid, uid)
await UniMessage.text(
f"了解~请发送图片吧!支持批量噢!\n如想退出搜索模式请发送“谢谢{NICKNAME}”"
).send()
await sleep(30)
ct = 0
while pls.get_on_off_status(gid):
if datetime.now() < pls.timeout[gid]:
if ct != pls.count[gid]:
ct = pls.count[gid]
pls.timeout[gid] = datetime.now() + timedelta(seconds=60)
else:
temp = pls.on[gid]
if not pls.count[gid]:
await UniMessage.at(user_id=str(temp)).text(
" 由于超时,已为您自动退出搜图模式~\n您本次搜索期间未发送任何图片,请检查是否被吞图~"
).send()
else:
await UniMessage.at(user_id=str(temp)).text(
f" 由于超时,已为您自动退出搜图模式,以后要记得说“谢谢{NICKNAME}”来退出搜图模式噢~\n您本次搜索共搜索了{pls.count[gid]}张图片~"
).send()
pls.turn_off(ev.group_id)
break
await sleep(30)
return
file, url, _ = ret
if str(uid) not in SUPERUSERS:
if not lmtd.check(uid):
await UniMessage.text(
f"您今天已经搜过{DAILY_LIMIT}次图了,休息一下明天再来吧~"
).send(at_sender=True)
return
if CHECK:
result = await check_screenshot(bot, file, url)
if result:
if result == 1:
await UniMessage.reply(id=str(mid)).text(
"该图似乎是手机截屏,请进行适当裁剪后再尝试搜图~\n*请注意搜索漫画时务必截取一个完整单页进行搜图~"
).send()
if result == 2:
await UniMessage.reply(id=str(mid)).text(
"该图似乎是长图拼接,请进行适当裁剪后再尝试搜图~\n*请注意搜索漫画时务必截取一个完整单页进行搜图~"
).send()
return
await UniMessage.text("正在搜索,请稍候~").send()
await picfinder(bot, ev, url)
@picmessage.handle()
async def picmessage_handle(bot: Bot, ev: GroupMessageEvent):
mid = ev.message_id
atcheck = False
batchcheck = False
for m in ev.message:
if m.type == "at" and int(m.data["qq"]) == ev.self_id:
atcheck = True
if pls.get_on_off_status(ev.group_id):
if int(pls.on[ev.group_id]) == int(ev.user_id):
batchcheck = True
if not (batchcheck or atcheck):
return
uid = ev.user_id
ret = _extract_image(ev)
if not ret:
return
file, url, sbtype = ret
if str(uid) not in SUPERUSERS:
if not lmtd.check(uid):
await UniMessage.text(
f"您今天已经搜过{DAILY_LIMIT}次图了,休息一下明天再来吧~"
).send(at_sender=True)
if pls.get_on_off_status(ev.group_id):
pls.turn_off(ev.group_id)
return
if pls.get_on_off_status(ev.group_id):
pls.count_plus(ev.group_id)
if pls.count[ev.group_id] > pls.limit[ev.group_id]:
await UniMessage.text(
f"您今天已经搜过{DAILY_LIMIT}次图了,休息一下明天再来吧~"
).send(at_sender=True)
pls.turn_off(ev.group_id)
return
if sbtype and IGNORE_STAMP:
if sbtype != "0":
await UniMessage.reply(id=str(mid)).text(
"该图为表情,已忽略~如确需搜索请尝试单发搜索或回复搜索~"
).send()
return
if CHECK:
result = await check_screenshot(bot, file, url)
if result:
if result == 1:
await UniMessage.reply(id=str(mid)).text(
"该图似乎是手机截屏,请手动进行适当裁剪后再尝试搜图~\n*请注意搜索漫画时务必截取一个完整单页进行搜图~"
).send()
if result == 2:
await UniMessage.reply(id=str(mid)).text(
"该图似乎是长图拼接,请手动进行适当裁剪后再尝试搜图~\n*请注意搜索漫画时务必截取一个完整单页进行搜图~"
).send()
return
if "c2cpicdw.qpic.cn/offpic_new/" in url:
md5 = file[:-6].upper()
url = f"http://gchat.qpic.cn/gchatpic_new/0/0-0-{md5}/0?term=2"
await UniMessage.text("正在搜索,请稍候~").send()
await picfinder(bot, ev, url)
@replymessage.handle()
async def replymessage_handle(bot: Bot, ev: GroupMessageEvent):
mid = ev.message_id
uid = ev.user_id
seg = ev.message[0]
if seg.type != "reply":
return
tmid = seg.data["id"]
cmd = ev.message.extract_plain_text()
flag1 = 0
flag2 = 0
for m in ev.message[2:]:
if m.type == "at" and int(m.data["qq"]) == ev.self_id:
flag1 = 1
for name in NICKNAMES:
if name in cmd:
flag1 = 1
break
for pfcmd in ["识图", "搜图", "查图", "找图"]:
if pfcmd in cmd:
flag2 = 1
if not (flag1 and flag2):
return
if str(uid) not in SUPERUSERS:
if not lmtd.check(uid):
await UniMessage.text(
f"您今天已经搜过{DAILY_LIMIT}次图了,休息一下明天再来吧~"
).send(at_sender=True)
try:
tmsg = await bot.get_msg(message_id=int(tmid))
except ActionFailed:
await UniMessage.text("该消息已过期,请重新转发~").send()
await replymessage.finish()
ret = re.search(r"\[CQ:image,file=(.*)?,url=(.*)\]", str(tmsg["message"]))
if not ret:
await UniMessage.text("未找到图片~").send()
return
file = ret.group(1)
url = ret.group(2)
if ",subType=" in url:
url = url.split(",")[0]
elif ",subType=" in file:
file = file.split(",")[0]
if CHECK:
result = await check_screenshot(bot, file, url)
if result:
if result == 1:
await UniMessage.reply(id=str(mid)).text(
"该图似乎是手机截屏,请手动进行适当裁剪后再尝试搜图~\n*请注意搜索漫画时务必截取一个完整单页进行搜图~"
).send()
if result == 2:
await UniMessage.reply(id=str(mid)).text(
"该图似乎是长图拼接,请手动进行适当裁剪后再尝试搜图~\n*请注意搜索漫画时务必截取一个完整单页进行搜图~"
).send()
return
if "c2cpicdw.qpic.cn/offpic_new/" in url:
md5 = file[:-6].upper()
url = f"http://gchat.qpic.cn/gchatpic_new/0/0-0-{md5}/0?term=2"
await UniMessage.text("正在搜索,请稍候~").send()
await picfinder(bot, ev, url)
@thanks.handle()
async def thanks_handle(bot: Bot, ev: GroupMessageEvent):
gid = ev.group_id
name = ev.message.extract_plain_text().removeprefix("谢谢").strip()
if name not in NICKNAMES:
return
if pls.get_on_off_status(gid):
if pls.on[gid] != ev.user_id:
await UniMessage.text("不能替别人结束搜图哦~").send()
return
if not pls.count[gid]:
await UniMessage.text("不用谢~\n您本次搜索期间未发送任何图片,请检查是否被吞图~").send()
else:
await UniMessage.text(f"不用谢~\n您本次搜索共搜索了{pls.count[gid]}张图片~").send()
pls.turn_off(gid)
return
await UniMessage.text("にゃ~").send()
@pic_private.handle()
async def pic_private_handle(bot: Bot, ev: PrivateMessageEvent):
uid = ev.user_id
ret = _extract_image(ev)
if not ret:
flag1 = flag2 = 0
for name in NICKNAMES:
if name in str(ev.message):
flag1 = 1
break
for pfcmd in ["识图", "搜图", "查图", "找图"]:
if pfcmd in str(ev.message):
flag2 = 1
if flag1 and flag2:
await UniMessage.text("私聊搜图请直接发送图片~").send()
return
file, url, _ = ret
if not lmtd.check(uid):
await UniMessage.text(
f"您今天已经搜过{DAILY_LIMIT}次图了,休息一下明天再来吧~"
).send()
return
if "c2cpicdw.qpic.cn/offpic_new/" in url:
md5 = file.upper()
logger.info(f"图片URL{url}")
url = f"https://gchat.qpic.cn/gchatpic_new/0/0-0-{md5}/0?term=2"
await UniMessage.text("正在搜索,请稍候~").send()
result = await get_image_data_sauce(url, SAUCENAO_KEY)
image_data_report = result[0]
simimax = result[1]
if "Index #" in image_data_report:
su = next(iter(SUPERUSERS))
await bot.send_private_msg(user_id=int(su), message="发生index解析错误")
await bot.send_private_msg(user_id=int(su), message=url)
await bot.send_private_msg(user_id=int(su), message=image_data_report)
await parse_report(image_data_report).send()
if float(simimax) > float(threshold):
lmtd.increase(uid)
else:
if simimax != 0:
await UniMessage.text("相似度过低,换用ascii2d检索中…").send()
else:
logger.error("SauceNao not found imageInfo")
await UniMessage.text("SauceNao检索失败,换用ascii2d检索中…").send()
image_data_report = await get_image_data_ascii(url)
if image_data_report[0]:
await parse_report(image_data_report[0]).send()
lmtd.increase(uid)
if image_data_report[1]:
await parse_report(image_data_report[1]).send()
if not (image_data_report[0] or image_data_report[1]):
logger.error("ascii2d not found imageInfo")
await UniMessage.text("ascii2d检索失败…").send()
async def chain_reply(bot: Bot, ev: GroupMessageEvent, chain, msg):
if not CHAIN_REPLY:
await parse_report(msg).send()
return chain
data = {
"type": "node",
"data": {
"name": str(NICKNAME) if str(NICKNAME) else "竹竹",
"uin": str(ev.self_id),
"content": str(msg),
},
}
chain.append(data)
return chain
async def picfinder(bot: Bot, ev: GroupMessageEvent, image_data):
uid = ev.user_id
chain = []
result = await get_image_data_sauce(image_data, SAUCENAO_KEY)
image_data_report = result[0]
simimax = result[1]
if "Index #" in image_data_report:
su = next(iter(SUPERUSERS))
await bot.send_private_msg(user_id=int(su), message="发生index解析错误")
await bot.send_private_msg(user_id=int(su), message=image_data)
await bot.send_private_msg(user_id=int(su), message=image_data_report)
chain = await chain_reply(bot, ev, chain, image_data_report)
if float(simimax) > float(threshold):
lmtd.increase(uid)
else:
if simimax != 0:
chain = await chain_reply(bot, ev, chain, "相似度过低,换用ascii2d检索中…")
else:
logger.error("SauceNao not found imageInfo")
chain = await chain_reply(bot, ev, chain, "SauceNao检索失败,换用ascii2d检索中…")
image_data_report = await get_image_data_ascii(image_data)
if image_data_report[0]:
chain = await chain_reply(bot, ev, chain, image_data_report[0])
lmtd.increase(uid)
if image_data_report[1]:
chain = await chain_reply(bot, ev, chain, image_data_report[1])
if not (image_data_report[0] or image_data_report[1]):
logger.error("ascii2d not found imageInfo")
chain = await chain_reply(bot, ev, chain, "ascii2d检索失败…")
if CHAIN_REPLY:
await bot.send_group_forward_msg(group_id=ev.group_id, messages=chain)
@@ -0,0 +1,50 @@
import os
threshold = 70 # SauceNAO相似度阈值,低于该相似度自动追加ascii2d搜索
SAUCENAO_KEY = "dfc2b9156f254754dee2ead531826020f6316f92" # SauceNAO 的 API key
SAUCENAO_RESULT_NUM = 3 # SauceNAO搜索结果显示数量
ASCII_RESULT_NUM = 3 # ascii2d搜索结果显示数量
SEARCH_TIMEOUT = 60 # 连续搜索模式超时时长
DAILY_LIMIT = 50 # 搜图每日限额
CHAIN_REPLY = True # 是否启用合并转发回复模式
THUMB_ON = True # 是否启用缩略图
CHECK = True # 是否开启手机截屏判定
IGNORE_STAMP = True # 是否在批量搜索中忽略表情包
HOST_CUSTOM = {
# 自定义Host,不使用留空即可
# 格式示例:'https://ascii2d.net' , 'http://localhost:12345'
'SAUCENAO': '',
'ASCII': ''
}
# ascii2d 的 Cloudflare cookie 文件(Netscape 格式,浏览器导出,参考 nonebot_plugin_video_analysis 的 cookies.txt 用法)
# 获取方式:浏览器打开 ascii2d.net 过验证后,用 EditThisCookie / Cookie-Editor 扩展导出为 Netscape 格式,
# 保存为 nonebot_plugin_picfinder_take/data/cookies_ascii2d.txt
# 留空则使用 playwright 无 cookie 尝试(可能无法通过 Cloudflare 验证)
ASCII2D_COOKIES_FILE = os.path.join(
os.path.dirname(os.path.abspath(__file__)), "data", "cookies_ascii2d.txt"
)
proxies = {
'http': '',
'https': ''
} # 网络代理
enableguild = {0000: [00000, 1111], 123456: [2345]} # 频道白名单 格式{频道id: [子频道id]}
helptext = '''
(xx为bot名称)
[@bot+图片] 单张/多张搜图
[xx搜图] 进入批量搜图模式
[谢谢xx] 退出批量搜图模式
'''.strip()
@@ -0,0 +1,865 @@
import asyncio
import base64
import os
import re
from io import BytesIO
from random import randint
from traceback import format_exc
import cloudscraper
import httpx
from PIL import Image
from lxml import etree
from nonebot.log import logger
from playwright.async_api import async_playwright
from .config import (
ASCII2D_COOKIES_FILE,
ASCII_RESULT_NUM,
HOST_CUSTOM,
SAUCENAO_RESULT_NUM,
THUMB_ON,
proxies,
)
def pic2b64(pic) -> str:
"""PIL Image 转 base64:// 字符串"""
buf = BytesIO()
pic.save(buf, format="PNG")
return "base64://" + base64.b64encode(buf.getvalue()).decode()
def _img_cq(base64_str: str) -> str:
"""生成图片 CQ 码字符串(用于嵌入搜索报告文本)"""
return f"[CQ:image,file={base64_str}]"
def _make_client(timeout: float) -> httpx.AsyncClient:
"""构造 httpx 客户端(httpx 0.28 移除了 AsyncClient 的 proxies 参数,转为 mounts)"""
http_proxy = proxies.get("http")
https_proxy = proxies.get("https")
if http_proxy or https_proxy:
mounts = {}
if http_proxy:
mounts["http://"] = httpx.HTTPTransport(proxy=http_proxy)
if https_proxy:
mounts["https://"] = httpx.HTTPTransport(proxy=https_proxy)
return httpx.AsyncClient(mounts=mounts, timeout=timeout)
return httpx.AsyncClient(timeout=timeout)
async def get_pic(address):
async with _make_client(20) as client:
resp = await client.get(address)
return resp.content
def randcolor():
return (randint(0, 255), randint(0, 255), randint(0, 255))
def ats_pic(img):
if img.mode != "RGB":
img = img.convert("RGB")
width = img.size[0] - 1 # 长度
height = img.size[1] - 1 # 宽度
img.putpixel((0, 0), randcolor())
img.putpixel((0, height), randcolor())
img.putpixel((width, 0), randcolor())
img.putpixel((width, height), randcolor())
return img
async def check_screenshot(bot, file, imgurl):
async with _make_client(20) as client:
pichead = await client.head(imgurl)
if pichead.headers["Content-Type"] == "image/gif":
print("gif pic, not likely a screen shot")
return 0
try:
resp = await client.get(imgurl)
image = Image.open(BytesIO(resp.content))
except Exception:
print("download failed")
return 0
cord = image.size[0] / image.size[1]
height = image.size[1]
print(cord)
if cord > 0.565:
print("too short, not likely a screen shot")
return 0
if cord < 0.2:
print("too long, might be long screen shot")
return 2
print("size checked, next ocr")
try:
ocr_result = await bot.call_api(".ocr_image", image=file)
except Exception:
print("ocr failed")
return False
flag = 0
for result in ocr_result["texts"]:
key1 = re.search("[0-9]{1,2}:[0-9]{2}", result["text"]) # 时间
key2 = re.search("移动|联通|电信", result["text"])
key3 = re.search("4G|5G", result["text"])
key4 = re.search("[0-9]{1,2}%", result["text"]) # 电量
key5 = re.search("[0-9]{0,3}[\\\/][0-9]{0,3}", result["text"]) # 页数
if key2 or key3 or key4:
print(str(result))
loc = result["coordinates"][2]["y"]
if int(loc) < (int(height) / 19):
flag = 1
if key1 or key5:
print(str(result))
loc = result["coordinates"][2]["y"]
if int(loc) < (int(height) / 19) or int(loc) > (int(height) * 18 / 20):
flag = 1
if flag:
break
if flag:
return 1
else:
return 0
def sauces_info(sauce):
service_name = ""
info = ""
try:
if sauce["header"]["index_id"] == 0:
service_name = "H-Magazines"
title = sauce["data"]["title"]
part = sauce["data"]["part"]
date = sauce["data"]["date"]
info = f"{title}-{part}/{date}"
# index 1 "h-anime" disabled
elif sauce["header"]["index_id"] == 2:
service_name = "H-Game CG"
company = sauce["data"]["company"]
title = sauce["data"]["title"]
info = f"[{company}] {title}"
# index 3 "ddb-objects" disabled
# index 4 "ddb-samples" disabled
elif sauce["header"]["index_id"] == 5 or sauce["header"]["index_id"] == 6:
service_name = "pixiv"
author_name = sauce["data"]["member_name"]
title = sauce["data"]["title"]
info = f"「{title}」/「{author_name}」"
# index 6 "pixiv historical" with 5
# index 7 "anime" disabled
elif sauce["header"]["index_id"] == 8:
service_name = "nico nico seiga"
author_name = sauce["data"]["member_name"]
title = sauce["data"]["title"]
info = f"「{title}」/「{author_name}」"
elif sauce["header"]["index_id"] == 9:
service_name = "Danbooru"
creator = sauce["data"]["creator"]
material = sauce["data"]["material"]
info = f"[{creator}]({material})"
elif sauce["header"]["index_id"] == 10:
service_name = "drawr Images"
author_name = sauce["data"]["member_name"]
title = sauce["data"]["title"]
info = f"「{title}」/「{author_name}」"
elif sauce["header"]["index_id"] == 11:
service_name = "Nijie Images"
author_name = sauce["data"]["member_name"]
title = sauce["data"]["title"]
info = f"「{title}」/「{author_name}」"
elif sauce["header"]["index_id"] == 12:
service_name = "Yande.re"
creator = sauce["data"]["creator"]
material = sauce["data"]["material"]
info = f"[{creator}]({material})"
# index 13 "animeop" disabled
# index 14 "IMDb" disabled
# index 15 "Shutterstock" disabled
elif sauce["header"]["index_id"] == 16:
service_name = "FAKKU"
creator = sauce["data"]["creator"]
source = sauce["data"]["source"]
info = f"[{creator}]({source})"
# index 17 reserved
elif sauce["header"]["index_id"] == 18 or sauce["header"]["index_id"] == 38:
service_name = "H-Misc (ehentai)"
eng_name = sauce["data"]["eng_name"]
jp_name = sauce["data"]["jp_name"]
info = f"{jp_name}" if jp_name else f"{eng_name}"
elif sauce["header"]["index_id"] == 19:
service_name = "2D-Market"
creator = sauce["data"]["creator"]
source = sauce["data"]["source"]
info = f"[{creator}]({source})"
elif sauce["header"]["index_id"] == 20:
service_name = "MediBang"
member_name = sauce["data"]["member_name"]
title = sauce["data"]["title"]
info = f"「{title}」/「{member_name}」"
elif sauce["header"]["index_id"] == 21:
service_name = "Anime"
title = sauce["data"]["source"]
year = sauce["data"]["year"]
part = sauce["data"]["part"]
est_time = sauce["data"]["est_time"]
time = est_time.split("/")[0]
info = f"《{title}》/{year}\n第{part}集,{time}"
elif sauce["header"]["index_id"] == 22:
service_name = "H-Anime"
title = sauce["data"]["source"]
year = sauce["data"]["year"]
part = sauce["data"]["part"]
est_time = sauce["data"]["est_time"]
time = est_time.split("/")[0]
info = f"《{title}》/{year}\n第{part}集,{time}"
elif sauce["header"]["index_id"] == 23:
service_name = "IMDb-Movies"
title = sauce["data"]["source"]
year = sauce["data"]["year"]
est_time = sauce["data"]["est_time"]
time = est_time.split("/")[0]
info = f"《{title}》/{year},{time}"
elif sauce["header"]["index_id"] == 24:
service_name = "IMDb-Shows"
title = sauce["data"]["source"]
year = sauce["data"]["year"]
part = sauce["data"]["part"]
est_time = sauce["data"]["est_time"]
time = est_time.split("/")[0]
info = f"《{title}》/{year}\n第{part}集,{time}"
elif sauce["header"]["index_id"] == 25:
service_name = "Gelbooru"
creator = sauce["data"]["creator"]
material = sauce["data"]["material"]
info = f"[{creator}]({material})"
elif sauce["header"]["index_id"] == 26:
service_name = "Konachan"
creator = sauce["data"]["creator"]
material = sauce["data"]["material"]
info = f"[{creator}]({material})"
elif sauce["header"]["index_id"] == 27:
service_name = "Sankaku Channel"
creator = sauce["data"]["creator"]
material = sauce["data"]["material"]
info = f"[{creator}]({material})"
elif sauce["header"]["index_id"] == 28:
service_name = "Anime-Pictures.net"
creator = sauce["data"]["creator"]
material = sauce["data"]["material"]
info = f"[{creator}]({material})"
elif sauce["header"]["index_id"] == 29:
service_name = "e621.net"
creator = sauce["data"]["creator"]
material = sauce["data"]["material"]
info = f"[{creator}]({material})"
elif sauce["header"]["index_id"] == 30:
service_name = "Idol Complex"
creator = sauce["data"]["creator"]
material = sauce["data"]["material"]
info = f"[{creator}]({material})"
elif sauce["header"]["index_id"] == 31:
service_name = "bcy.net Illust"
author_name = sauce["data"]["member_name"]
title = sauce["data"]["title"]
info = f"「{title}」/「{author_name}」"
elif sauce["header"]["index_id"] == 32:
service_name = "bcy.net Cosplay"
author_name = sauce["data"]["member_name"]
title = sauce["data"]["title"]
info = f"「{title}」/「{author_name}」"
elif sauce["header"]["index_id"] == 33:
service_name = "PortalGraphics.net"
member_name = sauce["data"]["member_name"]
title = sauce["data"]["title"]
info = f"「{title}」/「{member_name}」"
elif sauce["header"]["index_id"] == 34:
service_name = "deviantArt"
author_name = sauce["data"]["author_name"]
title = sauce["data"]["title"]
info = f"「{title}」/「{author_name}」"
elif sauce["header"]["index_id"] == 35:
service_name = "Pawoo.net"
illust_id = sauce["data"]["pawoo_id"]
author_name = sauce["data"]["pawoo_user_display_name"]
info = f"「{illust_id}」/「{author_name}」"
elif sauce["header"]["index_id"] == 36:
service_name = "Madokami (Manga)"
source = sauce["data"]["source"]
part = sauce["data"]["part"]
info = part if source in part else f"{source}-{part}"
elif sauce["header"]["index_id"] == 37 or sauce["header"]["index_id"] == 371:
service_name = "MangaDex"
artist = sauce["data"]["artist"]
author = sauce["data"]["author"]
source = sauce["data"]["source"]
part = sauce["data"]["part"]
info_a = f"[{artist}]" if artist == author else f"[{artist}({author})]"
info_b = part if source in part else f"{source}-{part}"
info = info_a + info_b
# index 38 "H-Misc (ehentai)" with 18
elif sauce["header"]["index_id"] == 39:
service_name = "Artstation"
author_name = sauce["data"]["author_name"]
title = sauce["data"]["title"]
info = f"「{title}」/「{author_name}」"
elif sauce["header"]["index_id"] == 40:
service_name = "FurAffinity"
author_name = sauce["data"]["author_name"]
title = sauce["data"]["title"]
info = f"「{title}」/「{author_name}」"
elif sauce["header"]["index_id"] == 41:
service_name = "Twitter"
author_name = sauce["data"]["twitter_user_handle"]
time = sauce["data"]["created_at"]
info = f"「{time[0:10]}」/「{author_name}」"
elif sauce["header"]["index_id"] == 42:
service_name = "Furry Network"
author_name = sauce["data"]["author_name"]
title = sauce["data"]["title"]
info = f"「{title}」/「{author_name}」"
elif sauce["header"]["index_id"] == 43:
service_name = "Kemono"
service = sauce["data"]["service_name"]
author_name = sauce["data"]["user_name"]
title = sauce["data"]["title"]
info = f"「{title}」/「({service}){author_name}」"
elif sauce["header"]["index_id"] == 44:
service_name = "Skeb"
creator_name = sauce["data"]["creator_name"]
creator = sauce["data"]["creator"]
info = f"[{creator_name}]({creator})"
else:
index = sauce["header"]["index_id"]
service_name = f"Index #{index}"
info = "no info"
except Exception as e:
index = sauce["header"]["index_id"]
service_name = f"Index #{index}"
info = "no info"
print(format_exc())
return service_name, info
class SauceNAO:
def __init__(
self,
api_key,
output_type=2,
testmode=0,
dbmask=None,
dbmaski=None,
db=999,
numres=3,
shortlimit=20,
longlimit=300,
):
params = dict()
params["api_key"] = api_key
params["output_type"] = output_type
params["testmode"] = testmode
params["dbmask"] = dbmask
params["dbmaski"] = dbmaski
params["db"] = db
params["numres"] = numres
self.params = params
self.host = HOST_CUSTOM["SAUCENAO"] or "https://saucenao.com"
self.header = "————>saucenao<————"
async def get_sauce(self, image_url):
logger.debug(f"Now starting get the SauceNAO data:{image_url}")
# 过滤 None 参数(SauceNAO 文件上传遇 dbmask=None 编码会返回 500)
params = {k: v for k, v in self.params.items() if v is not None}
# 优先本地下载图片后以文件形式提交,避免 SauceNAO 服务器抓取 QQ 图床失败
try:
async with _make_client(20) as client:
resp = await client.get(image_url)
resp.raise_for_status()
image_bytes = resp.content
submit_ok = True
except Exception as e:
logger.error(f"图片本地下载失败({e}), 退回 URL 提交")
submit_ok = False
if submit_ok:
params.pop("url", None)
async with _make_client(30) as client:
response = await client.post(
f"{self.host}/search.php",
params=params,
files={"file": ("search_image.jpg", image_bytes, "image/jpeg")},
)
else:
params["url"] = image_url
async with _make_client(15) as client:
response = await client.get(f"{self.host}/search.php", params=params)
if response.status_code != 200:
raise RuntimeError(
f"SauceNAO 请求失败: HTTP {response.status_code}, 响应: {response.text[:100]!r}"
)
return response.json()
async def get_view(self, sauce) -> str:
sauces = await self.get_sauce(sauce)
repass = ""
simimax = 0
index = 1
for sauce in sauces["results"]:
try:
url = (
sauce["data"]["ext_urls"][0].replace("\\", "").strip()
if "ext_urls" in sauce["data"]
else "no link"
)
similarity = sauce["header"]["similarity"]
if not similarity.replace(".", "").isdigit():
similarity = 0
simimax = float(similarity) if float(similarity) > simimax else simimax
thumbnail_url = sauce["header"]["thumbnail"]
if THUMB_ON:
try:
thumbnail_image = _img_cq(
pic2b64(
ats_pic(
Image.open(BytesIO(await get_pic(thumbnail_url)))
)
)
)
except Exception as e:
print(format_exc())
thumbnail_image = "[预览图下载失败]"
else:
thumbnail_image = ""
service_name, info = sauces_info(sauce)
putline = (
f"搜索结果{index}:\n\n{thumbnail_image}\n\n平台:\n\n{service_name}:\n\n"
f"相似度:{similarity}%\n\n信息:\n\n{info}\n\n相关:\n\n{url}"
)
if repass:
repass = "\n\n\n".join([repass, putline])
else:
repass = putline
except Exception as e:
print(format_exc())
pass
index += 1
return [repass, simimax]
headers = {
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9",
"Accept-Encoding": "gzip, deflate",
"Accept-Language": "zh-CN,zh;q=0.9",
"Cache-Control": "max-age=0",
"Connection": "keep-alive",
"Origin": "https://ascii2d.net",
"Referer": "https://ascii2d.net/",
"Sec-Fetch-Dest": "document",
"Sec-Fetch-Mode": "navigate",
"Sec-Fetch-Site": "same-origin",
"Sec-Fetch-User": "?1",
"Upgrade-Insecure-Requests": "1",
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.163 Safari/537.36",
}
class ascii2d:
def __init__(self, num=2):
self.num = num
self.host = HOST_CUSTOM["ASCII"] or "https://ascii2d.net"
self.header = "————>ascii2d<————"
self.scraper = cloudscraper.create_scraper()
async def get_search_data(self, url: str, data=None):
if data is not None:
html = data
else:
html = await get_html_ascii2d(url)
all_data = html.xpath('//div[@class="row item-box"]')
info = []
for data in all_data[1 : self.num + 1]:
try:
title = ""
member = ""
if not data.xpath('.//img[@loading="lazy"]/@src'):
continue
thumb_url = data.xpath('.//img[@loading="lazy"]/@src')[0].strip()
thumb_url = f"{self.host}{thumb_url}"
if not data.xpath('.//div[@class="detail-box gray-link"]/h6'):
data2 = (
data.xpath('.//div[@class="external"]')[0]
if data.xpath('.//div[@class="external"]')
else data
)
info_url = (
data2.xpath(".//a/@href")[0].strip()
if data2.xpath(".//a/@rel")
else "no link"
)
tag = "外部登录" if info_url == "no link" else info_url.split("/")[2]
else:
data2 = data.xpath('.//div[@class="detail-box gray-link"]/h6')[0]
info_url = data2.xpath(".//a/@href")[0].strip()
tag = (
data2.xpath("./small/text()") or data2.xpath(".//a/text()")
)[0].strip()
if tag == "pixiv" or tag == "twitter":
title = data2.xpath(".//a//text()")[0]
member = data2.xpath(".//a//text()")[1]
title = f"「{title}」/「{member}」"
elif tag == "外部登录":
title = data2.text.replace("\n", "") if data2.text else ""
else:
title = data2.text.replace("\n", "")
title = f"「{title}」"
info.append([info_url, tag, thumb_url, title])
except Exception as e:
print(format_exc())
logger.error(e)
continue
return info
async def add_repass(self, tag: str, data, thumbnails: dict | None = None):
po = "——{}——".format(tag)
index = 1
for line in data:
if THUMB_ON:
try:
thumb_content = (thumbnails or {}).get(line[2])
if thumb_content is None:
thumb_content = self.scraper.get(
line[2], timeout=20, proxies=proxies
).content
thumbnail_image = _img_cq(
pic2b64(ats_pic(Image.open(BytesIO(thumb_content))))
)
except Exception as e:
print(format_exc())
thumbnail_image = "[预览图下载失败]"
else:
thumbnail_image = ""
putline = (
f"搜索结果{index}:\n\n{thumbnail_image}\n\n平台:\n\n{line[1]}\n\n"
f"信息:\n\n{line[3]}\n\n相关:\n\n{line[0]}"
)
po = "\n\n\n".join([po, putline])
index += 1
return po
async def get_view(self, ascii2d) -> str:
putline1 = ""
putline2 = ""
url_index = f"{self.host}/search/url/{ascii2d}"
logger.debug(f"Now starting get the {url_index}")
try:
html_index = await get_html_ascii2d(url_index)
if html_index is None:
logger.error("ascii2d 页面获取失败(Cloudflare challenge 未通过)")
return [putline1, putline2]
except Exception as e:
print(format_exc())
logger.error(f"ascii2d get html data failed: {e}")
return [putline1, putline2]
neet_div = html_index.xpath(
'//div[@class="detail-link pull-xs-right hidden-sm-down gray-link"]'
)
if neet_div:
a_url_foot = neet_div[0].xpath("./span/a/@href")
url2 = f"{self.host}{a_url_foot[1]}"
color = await self.get_search_data("", data=html_index)
bovw = await self.get_search_data(url2)
# 用系统浏览器会话批量下载缩略图(带 CF cookie,避免被 Cloudflare 拦截)
thumb_urls = [line[2] for line in color + bovw]
thumbnails = await fetch_thumbs_playwright(thumb_urls)
if color:
putline1 = await self.add_repass("色调检索", color, thumbnails)
if bovw:
putline2 = await self.add_repass("特征检索", bovw, thumbnails)
return [putline1, putline2]
async def get_image_data_sauce(image_url: str, api_key: str):
if type(image_url) == list:
image_url = image_url[0]
logger.info("Loading Image Search Container……")
NAO = SauceNAO(api_key, numres=SAUCENAO_RESULT_NUM)
logger.debug("Loading all view……")
repass = ""
simimax = 0
# 网络抖动/免费key限流时重试,退避间隔同时规避 SauceNAO 的 4 秒限流
for attempt in range(3):
try:
result = await NAO.get_view(image_url)
if result:
header = NAO.header
simimax = result[1]
repass = "\n".join([header, result[0]])
break
except Exception as e:
logger.error(f"SauceNAO 第{attempt + 1}次尝试失败: {e}")
if attempt < 2:
await asyncio.sleep(2 * (attempt + 1))
else:
return ["SauceNAO搜索失败……", 0]
return [repass, simimax]
async def get_image_data_ascii(image_url: str):
if type(image_url) == list:
image_url = image_url[0]
logger.info("Loading Image Search Container……")
ii2d = ascii2d(ASCII_RESULT_NUM)
logger.debug("Loading all view……")
repass1 = ""
repass2 = ""
try:
putline = await ii2d.get_view(image_url)
if putline:
header = ii2d.header
if putline[0]:
repass1 = "\n".join([header, putline[0]])
if putline[1]:
repass2 = "\n".join([header, putline[1]])
except Exception as e:
logger.error(format_exc())
return ["ascii2d搜索失败……", ""]
return [repass1, repass2]
# 系统浏览器可执行文件(与获取 cookie 的浏览器同款 TLS 指纹)
CHROME_PATHS = [
"C:/Program Files/Google/Chrome/Application/chrome.exe",
"C:/Program Files (x86)/Google/Chrome/Application/chrome.exe",
"C:/Program Files (x86)/Microsoft/Edge/Application/msedge.exe",
"C:/Program Files/Microsoft/Edge/Application/msedge.exe",
]
# 专用浏览器 profile 目录(cf_clearance 持久化,首次过验证后长期有效)
CHROME_PROFILE_DIR = os.path.join(
os.path.dirname(os.path.abspath(__file__)), "data", "chrome_profile"
)
def _parse_netscape_cookies(file_path: str) -> list[dict]:
"""解析 Netscape 格式 cookies 文件(与 nonebot_plugin_video_analysis 同款)"""
cookies = []
with open(file_path, "r", encoding="utf-8") as f:
for line in f:
line = line.strip()
if not line or line.startswith("#"):
continue
parts = line.split("\t")
if len(parts) != 7:
continue
domain, flag, path, secure, expiry, name, value = parts
cookie = {
"name": name,
"value": value,
"domain": domain,
"path": path,
"secure": secure.upper() == "TRUE",
"sameSite": "Lax",
}
if expiry.isdigit() and int(expiry) > 0:
cookie["expires"] = int(expiry)
cookies.append(cookie)
return cookies
def _load_ascii2d_cookies():
"""读取配置的 ascii2d cookie 文件,未配置或不存在返回 None"""
if not ASCII2D_COOKIES_FILE or not os.path.isfile(ASCII2D_COOKIES_FILE):
return None
return _parse_netscape_cookies(ASCII2D_COOKIES_FILE)
def _save_netscape_cookies(cookies: list[dict], file_path: str):
"""把浏览器回传的 cookies 写回 Netscape 格式文件(实现 cf_clearance 自愈)"""
lines = ["# Netscape HTTP Cookie File"]
for c in cookies:
domain = c.get("domain", "")
include_subdomains = "TRUE" if domain.startswith(".") else "FALSE"
path = c.get("path", "/")
secure = "TRUE" if c.get("secure") else "FALSE"
try:
expires = int(c.get("expires") or 0)
except (TypeError, ValueError):
expires = 0
name = c.get("name", "")
value = c.get("value", "")
lines.append(
f"{domain}\t{include_subdomains}\t{path}\t{secure}\t{expires}\t{name}\t{value}"
)
with open(file_path, "w", encoding="utf-8") as f:
f.write("\n".join(lines))
async def get_html_ascii2d(url):
"""获取 ascii2d 页面 HTML:配置了 cookie 文件时注入系统浏览器直过验证"""
return await get_html_content(url, cookies=_load_ascii2d_cookies())
async def fetch_thumbs_playwright(urls: list[str]) -> dict:
"""用系统浏览器会话批量下载缩略图(带 CF cookie,避免 Cloudflare 拦截)"""
if not urls:
return {}
result = {}
async with async_playwright() as p:
launch_kwargs = {
"headless": False,
"args": [
"--window-position=-32000,-32000",
"--disable-blink-features=AutomationControlled",
],
}
for path in CHROME_PATHS:
if os.path.isfile(path):
launch_kwargs["executable_path"] = path
break
context = await p.chromium.launch_persistent_context(
CHROME_PROFILE_DIR, **launch_kwargs
)
cookies = _load_ascii2d_cookies()
if cookies:
await context.add_cookies(cookies)
try:
for u in urls:
try:
resp = await context.request.get(u, timeout=20000)
if resp.status == 200 and "image" in resp.headers.get(
"content-type", ""
):
result[u] = await resp.body()
except Exception as e:
logger.debug(f"缩略图下载失败 {u}: {e}")
finally:
await context.close()
return result
async def get_html_content(url, cookies=None):
async with async_playwright() as p:
# 用系统 Chrome + 固定 profile 目录(cf_clearance 持久化,与浏览器获取 cookie 时同款 TLS 指纹)
# 不指定 UA:使用系统 Chrome 真实 UA,否则 cf_clearance(与 UA 绑定)会失效
# headless 会被 Cloudflare 检测拒绝,用有头模式并将窗口移出屏幕避免打扰
launch_kwargs = {
"headless": False,
"args": [
"--window-position=-32000,-32000",
"--disable-blink-features=AutomationControlled",
],
}
for path in CHROME_PATHS:
if os.path.isfile(path):
launch_kwargs["executable_path"] = path
break
context = await p.chromium.launch_persistent_context(
CHROME_PROFILE_DIR, **launch_kwargs
)
await context.add_init_script(
"Object.defineProperty(navigator, 'webdriver', {get: () => undefined});"
)
if cookies:
await context.add_cookies(cookies)
page = context.pages[0] if context.pages else await context.new_page()
try:
await page.goto(url, wait_until="domcontentloaded")
# 轮询等待 Cloudflare challenge 自动通过(最多约 20 秒)
passed = False
for _ in range(8):
await page.wait_for_timeout(2500)
title = await page.title()
if "Just a moment" not in title:
passed = True
break
html = etree.HTML(await page.content())
# 浏览器可能刷新了 cf_clearance,回写文件实现自愈
if passed and ASCII2D_COOKIES_FILE:
try:
_save_netscape_cookies(
await context.cookies(), ASCII2D_COOKIES_FILE
)
logger.info("ascii2d cookies 已刷新回写")
except Exception as e:
logger.error(f"ascii2d cookies 回写失败: {e}")
print(f"获取到的html内容:{html}")
return html
except Exception as e:
print(f"Error fetching page: {e}")
return None
finally:
await context.close()