<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>DrissionPage on heyaohua's Blog</title><link>https://blog.heyaohua.com/tags/drissionpage/</link><description>Recent content in DrissionPage on heyaohua's Blog</description><image><title>heyaohua's Blog</title><url>https://blog.heyaohua.com/og-image.png</url><link>https://blog.heyaohua.com/og-image.png</link></image><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Fri, 26 Sep 2025 14:00:00 +0800</lastBuildDate><atom:link href="https://blog.heyaohua.com/tags/drissionpage/index.xml" rel="self" type="application/rss+xml"/><item><title>淘宝自动化框架选择方案</title><link>https://blog.heyaohua.com/posts/2025/09/taobao-automation-framework/</link><pubDate>Fri, 26 Sep 2025 14:00:00 +0800</pubDate><guid>https://blog.heyaohua.com/posts/2025/09/taobao-automation-framework/</guid><description>国产框架，中文文档完善</description><content:encoded><![CDATA[<h1 id="淘宝自动化框架选择方案">淘宝自动化框架选择方案</h1>
<h2 id="-推荐方案drissionpage--现有架构">🎯 推荐方案：DrissionPage + 现有架构</h2>
<h3 id="为什么选择-drissionpage">为什么选择 DrissionPage？</h3>
<ol>
<li><strong>专为中国网站设计</strong></li>
<li>针对淘宝、京东等电商网站优化</li>
<li>内置常见反爬虫机制绕过</li>
<li></li>
</ol>
<p>国产框架，中文文档完善</p>
<ol start="5">
<li></li>
</ol>
<p><strong>与现有架构完美融合</strong></p>
<ol start="6">
<li>可以直接使用现有的 requests session</li>
<li>支持与 mitmproxy 代理集成</li>
<li></li>
</ol>
<p>兼容现有的数据处理管道</p>
<ol start="9">
<li></li>
</ol>
<p><strong>性能与易用性并存</strong></p>
<ol start="10">
<li>基于 Chromium 内核，性能优秀</li>
<li>API 设计简洁直观</li>
<li>支持页面模式和 requests 模式切换</li>
</ol>
<h2 id="-框架对比分析">📊 框架对比分析</h2>
<table>
  <thead>
      <tr>
          <th>特性</th>
          <th>DrissionPage</th>
          <th>Playwright</th>
          <th>Selenium</th>
          <th>Requests-HTML</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><strong>性能</strong></td>
          <td>很快</td>
          <td>最快</td>
          <td>中等</td>
          <td>快</td>
      </tr>
      <tr>
          <td><strong>反爬虫能力</strong></td>
          <td>优秀</td>
          <td>优秀</td>
          <td>一般</td>
          <td>较弱</td>
      </tr>
      <tr>
          <td><strong>淘宝适配</strong></td>
          <td>优秀</td>
          <td>好</td>
          <td>一般</td>
          <td>较弱</td>
      </tr>
      <tr>
          <td><strong>学习成本</strong></td>
          <td>低</td>
          <td>中</td>
          <td>中</td>
          <td>低</td>
      </tr>
      <tr>
          <td><strong>中文文档</strong></td>
          <td>优秀</td>
          <td>一般</td>
          <td>好</td>
          <td>一般</td>
      </tr>
      <tr>
          <td><strong>社区支持</strong></td>
          <td>活跃</td>
          <td>活跃</td>
          <td>最大</td>
          <td>较小</td>
      </tr>
  </tbody>
</table>
<h2 id="-技术实施路线">🛠️ 技术实施路线</h2>
<h3 id="阶段一环境准备">阶段一：环境准备</h3>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#282a36;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"><code class="language-bash" data-lang="bash"><span style="display:flex;"><span><span style="color:#6272a4"># 安装 DrissionPage</span>
</span></span><span style="display:flex;"><span>pip install DrissionPage
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#6272a4"># 安装备选方案（可选）</span>
</span></span><span style="display:flex;"><span>pip install playwright
</span></span><span style="display:flex;"><span>pip install selenium
</span></span></code></pre></div><h3 id="阶段二基础集成">阶段二：基础集成</h3>
<ol>
<li>创建 <code>TaobaoAutomator</code> 类</li>
<li>集成现有的代理服务器</li>
<li>实现基础的搜索和数据提取功能</li>
</ol>
<h3 id="阶段三高级功能">阶段三：高级功能</h3>
<ol>
<li>反爬虫策略优化</li>
<li>数据清洗和存储</li>
<li>错误处理和重试机制</li>
</ol>
<h3 id="阶段四性能优化">阶段四：性能优化</h3>
<ol>
<li>并发处理</li>
<li>资源管理</li>
<li>监控和日志</li>
</ol>
<h2 id="-备选方案">💡 备选方案</h2>
<h3 id="方案-a纯-playwright如果团队技术能力强">方案 A：纯 Playwright（如果团队技术能力强）</h3>
<ul>
<li>性能最佳</li>
<li>功能最全面</li>
<li>需要较多学习时间</li>
</ul>
<h3 id="方案-bselenium如果需要最大兼容性">方案 B：Selenium（如果需要最大兼容性）</h3>
<ul>
<li>社区资源最丰富</li>
<li>兼容性最好</li>
<li>性能相对较慢</li>
</ul>
<h3 id="方案-c混合方案">方案 C：混合方案</h3>
<ul>
<li>DrissionPage 处理复杂交互</li>
<li>requests 处理简单API调用</li>
<li>mitmproxy 处理数据截取</li>
</ul>
<h2 id="-具体实现示例">🎪 具体实现示例</h2>
<h3 id="drissionpage-基础用法">DrissionPage 基础用法</h3>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#282a36;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#ff79c6">from</span> DrissionPage <span style="color:#ff79c6">import</span> ChromiumPage
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#6272a4"># 创建页面对象</span>
</span></span><span style="display:flex;"><span>page <span style="color:#ff79c6">=</span> ChromiumPage()
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#6272a4"># 访问淘宝</span>
</span></span><span style="display:flex;"><span>page<span style="color:#ff79c6">.</span>get(<span style="color:#f1fa8c">&#39;https://www.taobao.com&#39;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#6272a4"># 搜索商品</span>
</span></span><span style="display:flex;"><span>search_box <span style="color:#ff79c6">=</span> page<span style="color:#ff79c6">.</span>ele(<span style="color:#f1fa8c">&#39;#q&#39;</span>)
</span></span><span style="display:flex;"><span>search_box<span style="color:#ff79c6">.</span>input(<span style="color:#f1fa8c">&#39;手机&#39;</span>)
</span></span><span style="display:flex;"><span>search_box<span style="color:#ff79c6">.</span>after()<span style="color:#ff79c6">.</span>click()
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#6272a4"># 获取商品信息</span>
</span></span><span style="display:flex;"><span>products <span style="color:#ff79c6">=</span> page<span style="color:#ff79c6">.</span>eles(<span style="color:#f1fa8c">&#39;.item&#39;</span>)
</span></span><span style="display:flex;"><span><span style="color:#ff79c6">for</span> product <span style="color:#ff79c6">in</span> products:
</span></span><span style="display:flex;"><span>    title <span style="color:#ff79c6">=</span> product<span style="color:#ff79c6">.</span>ele(<span style="color:#f1fa8c">&#39;.title&#39;</span>)<span style="color:#ff79c6">.</span>text
</span></span><span style="display:flex;"><span>    price <span style="color:#ff79c6">=</span> product<span style="color:#ff79c6">.</span>ele(<span style="color:#f1fa8c">&#39;.price&#39;</span>)<span style="color:#ff79c6">.</span>text
</span></span><span style="display:flex;"><span>    <span style="color:#8be9fd;font-style:italic">print</span>(<span style="color:#f1fa8c">f</span><span style="color:#f1fa8c">&#34;</span><span style="color:#f1fa8c">{</span>title<span style="color:#f1fa8c">}</span><span style="color:#f1fa8c">: </span><span style="color:#f1fa8c">{</span>price<span style="color:#f1fa8c">}</span><span style="color:#f1fa8c">&#34;</span>)
</span></span></code></pre></div><h3 id="与现有架构集成">与现有架构集成</h3>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#282a36;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#ff79c6">from</span> DrissionPage <span style="color:#ff79c6">import</span> ChromiumPage
</span></span><span style="display:flex;"><span><span style="color:#ff79c6">from</span> crawler.gateway.proxy_server <span style="color:#ff79c6">import</span> ProxyServer
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#ff79c6">class</span> <span style="color:#50fa7b">TaobaoAutomator</span>:
</span></span><span style="display:flex;"><span>    <span style="color:#ff79c6">def</span> <span style="color:#50fa7b">__init__</span>(<span style="font-style:italic">self</span>):
</span></span><span style="display:flex;"><span>        <span style="color:#6272a4"># 启动代理服务器</span>
</span></span><span style="display:flex;"><span>        <span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>proxy_server <span style="color:#ff79c6">=</span> ProxyServer()
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>        <span style="color:#6272a4"># 配置 DrissionPage 使用代理</span>
</span></span><span style="display:flex;"><span>        <span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>page <span style="color:#ff79c6">=</span> ChromiumPage()
</span></span><span style="display:flex;"><span>        <span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>page<span style="color:#ff79c6">.</span>set<span style="color:#ff79c6">.</span>proxy(<span style="color:#f1fa8c">f</span><span style="color:#f1fa8c">&#39;127.0.0.1:</span><span style="color:#f1fa8c">{</span><span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>proxy_server<span style="color:#ff79c6">.</span>port<span style="color:#f1fa8c">}</span><span style="color:#f1fa8c">&#39;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#ff79c6">def</span> <span style="color:#50fa7b">search_products</span>(<span style="font-style:italic">self</span>, keyword):
</span></span><span style="display:flex;"><span>        <span style="color:#6272a4"># 实现搜索逻辑</span>
</span></span><span style="display:flex;"><span>        <span style="color:#ff79c6">pass</span>
</span></span></code></pre></div><h2 id="-技术要点">🔧 技术要点</h2>
<ol>
<li><strong>代理集成</strong>：确保自动化框架使用现有的代理服务器</li>
<li><strong>数据同步</strong>：截取的API数据与页面数据关联</li>
<li><strong>反爬虫</strong>：实现用户行为模拟和请求间隔控制</li>
<li><strong>错误处理</strong>：网络异常、页面变化等情况的处理</li>
</ol>
<h2 id="-预期效果">📈 预期效果</h2>
<ul>
<li><strong>开发效率提升 50%</strong>：相比从零开始</li>
<li><strong>数据质量提升</strong>：结合API和页面数据</li>
<li><strong>稳定性增强</strong>：多重反爬虫策略</li>
<li><strong>维护成本降低</strong>：统一的架构设计</li>
</ul>
]]></content:encoded></item><item><title>我用Python开发了一个淘宝图片搜索自动化系统</title><link>https://blog.heyaohua.com/posts/2025/05/taobao-image-search-automation/</link><pubDate>Mon, 26 May 2025 16:00:00 +0800</pubDate><guid>https://blog.heyaohua.com/posts/2025/05/taobao-image-search-automation/</guid><description>在电商时代，图片搜索已经成为用户发现商品的重要方式。作为开发者，我经常需要为客户批量搜索相似商品并生成报告。手动操作不仅效率低下，还容易出错。于是，我决定开发一个自动化系统来解决这个问题。</description><content:encoded><![CDATA[<p>在电商时代，图片搜索已经成为用户发现商品的重要方式。作为开发者，我经常需要为客户批量搜索相似商品并生成报告。手动操作不仅效率低下，还容易出错。于是，我决定开发一个自动化系统来解决这个问题。</p>
<h2 id="项目目标">项目目标</h2>
<ul>
<li>批量处理图片搜索</li>
<li>自动提取商品数据</li>
<li>生成包含图片的Excel报告</li>
<li>自动发送邮件通知</li>
<li>完整的错误处理和日志记录</li>
</ul>
<h2 id="技术选型">技术选型</h2>
<h3 id="自动化框架drissionpage">自动化框架：DrissionPage</h3>
<p>经过对比Selenium、Playwright等框架，我选择了DrissionPage：</p>
<ul>
<li>专为中国网站优化</li>
<li>反爬虫能力强</li>
<li>对淘宝等国内电商支持好</li>
</ul>
<h3 id="数据拦截mitmproxy">数据拦截：mitmproxy</h3>
<ul>
<li>能够拦截HTTPS流量</li>
<li>支持自定义插件</li>
<li>适合API数据提取</li>
</ul>
<h3 id="数据处理">数据处理</h3>
<ul>
<li>Pandas：数据处理</li>
<li>openpyxl：Excel操作</li>
<li>Pillow：图片处理</li>
</ul>
<h2 id="核心功能实现">核心功能实现</h2>
<h3 id="1-图片搜索自动化">1. 图片搜索自动化</h3>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#282a36;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#ff79c6">def</span> <span style="color:#50fa7b">search_by_image</span>(<span style="font-style:italic">self</span>, image_path: <span style="color:#8be9fd;font-style:italic">str</span>):
</span></span><span style="display:flex;"><span>    <span style="color:#f1fa8c">&#34;&#34;&#34;图片搜索功能&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    <span style="color:#6272a4"># 1. 打开淘宝首页</span>
</span></span><span style="display:flex;"><span>    <span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>browser<span style="color:#ff79c6">.</span>get(<span style="color:#f1fa8c">&#39;https://www.taobao.com&#39;</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#6272a4"># 2. 点击搜同款按钮</span>
</span></span><span style="display:flex;"><span>    search_button <span style="color:#ff79c6">=</span> <span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>browser<span style="color:#ff79c6">.</span>ele(<span style="color:#f1fa8c">&#39;css:.image-search-icon-wrapper&#39;</span>)
</span></span><span style="display:flex;"><span>    search_button<span style="color:#ff79c6">.</span>click()
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#6272a4"># 3. 上传图片</span>
</span></span><span style="display:flex;"><span>    file_input <span style="color:#ff79c6">=</span> <span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>browser<span style="color:#ff79c6">.</span>ele(<span style="color:#f1fa8c">&#39;css:#image-search-custom-file-input&#39;</span>)
</span></span><span style="display:flex;"><span>    file_input<span style="color:#ff79c6">.</span>input(image_path)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#6272a4"># 4. 等待上传完成并搜索</span>
</span></span><span style="display:flex;"><span>    <span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>_wait_for_upload_complete()
</span></span><span style="display:flex;"><span>    search_btn <span style="color:#ff79c6">=</span> <span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>browser<span style="color:#ff79c6">.</span>ele(<span style="color:#f1fa8c">&#39;css:#image-search-upload-button&#39;</span>)
</span></span><span style="display:flex;"><span>    search_btn<span style="color:#ff79c6">.</span>click()
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#6272a4"># 5. 提取商品数据</span>
</span></span><span style="display:flex;"><span>    <span style="color:#ff79c6">return</span> <span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>_extract_products_from_page()
</span></span></code></pre></div><h3 id="2-数据拦截与提取">2. 数据拦截与提取</h3>
<p>通过mitmproxy拦截淘宝API响应，提取商品信息：</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#282a36;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#ff79c6">def</span> <span style="color:#50fa7b">response</span>(flow: http<span style="color:#ff79c6">.</span>HTTPFlow) <span style="color:#ff79c6">-&gt;</span> <span style="color:#ff79c6">None</span>:
</span></span><span style="display:flex;"><span>    <span style="color:#f1fa8c">&#34;&#34;&#34;拦截API响应&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    <span style="color:#ff79c6">if</span> <span style="color:#f1fa8c">&#39;h5api.m.taobao.com&#39;</span> <span style="color:#ff79c6">in</span> flow<span style="color:#ff79c6">.</span>request<span style="color:#ff79c6">.</span>pretty_url:
</span></span><span style="display:flex;"><span>        content <span style="color:#ff79c6">=</span> flow<span style="color:#ff79c6">.</span>response<span style="color:#ff79c6">.</span>text
</span></span><span style="display:flex;"><span>        <span style="color:#6272a4"># 解析JSONP响应，提取商品数据</span>
</span></span><span style="display:flex;"><span>        data <span style="color:#ff79c6">=</span> parse_jsonp_response(content)
</span></span><span style="display:flex;"><span>        save_to_file(data)
</span></span></code></pre></div><h3 id="3-excel报告生成">3. Excel报告生成</h3>
<p>生成多Sheet的Excel文件，包含压缩图片：</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#282a36;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#ff79c6">def</span> <span style="color:#50fa7b">generate_excel_report</span>(<span style="font-style:italic">self</span>, products_data):
</span></span><span style="display:flex;"><span>    <span style="color:#f1fa8c">&#34;&#34;&#34;生成Excel报告&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    workbook <span style="color:#ff79c6">=</span> openpyxl<span style="color:#ff79c6">.</span>Workbook()
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#ff79c6">for</span> sheet_data <span style="color:#ff79c6">in</span> products_data:
</span></span><span style="display:flex;"><span>        worksheet <span style="color:#ff79c6">=</span> workbook<span style="color:#ff79c6">.</span>create_sheet(sheet_data[<span style="color:#f1fa8c">&#39;name&#39;</span>])
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>        <span style="color:#6272a4"># 添加商品数据</span>
</span></span><span style="display:flex;"><span>        <span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>add_product_data(worksheet, sheet_data[<span style="color:#f1fa8c">&#39;products&#39;</span>])
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>        <span style="color:#6272a4"># 下载并添加商品图片</span>
</span></span><span style="display:flex;"><span>        <span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>add_product_images(worksheet, sheet_data[<span style="color:#f1fa8c">&#39;products&#39;</span>])
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    workbook<span style="color:#ff79c6">.</span>save(<span style="color:#f1fa8c">&#39;report.xlsx&#39;</span>)
</span></span></code></pre></div><h3 id="4-图片压缩优化">4. 图片压缩优化</h3>
<p>解决Excel文件过大的问题：</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#282a36;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#ff79c6">def</span> <span style="color:#50fa7b">compress_image</span>(<span style="font-style:italic">self</span>, image_path: <span style="color:#8be9fd;font-style:italic">str</span>) <span style="color:#ff79c6">-&gt;</span> <span style="color:#8be9fd;font-style:italic">str</span>:
</span></span><span style="display:flex;"><span>    <span style="color:#f1fa8c">&#34;&#34;&#34;智能图片压缩&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    <span style="color:#ff79c6">with</span> Image<span style="color:#ff79c6">.</span>open(image_path) <span style="color:#ff79c6">as</span> img:
</span></span><span style="display:flex;"><span>        <span style="color:#6272a4"># 调整尺寸到400x400</span>
</span></span><span style="display:flex;"><span>        <span style="color:#ff79c6">if</span> img<span style="color:#ff79c6">.</span>width <span style="color:#ff79c6">&gt;</span> <span style="color:#bd93f9">400</span> <span style="color:#ff79c6">or</span> img<span style="color:#ff79c6">.</span>height <span style="color:#ff79c6">&gt;</span> <span style="color:#bd93f9">400</span>:
</span></span><span style="display:flex;"><span>            img<span style="color:#ff79c6">.</span>thumbnail((<span style="color:#bd93f9">400</span>, <span style="color:#bd93f9">400</span>), Image<span style="color:#ff79c6">.</span>Resampling<span style="color:#ff79c6">.</span>LANCZOS)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>        <span style="color:#6272a4"># 压缩质量到40%</span>
</span></span><span style="display:flex;"><span>        img<span style="color:#ff79c6">.</span>save(compressed_path, <span style="color:#f1fa8c">&#34;JPEG&#34;</span>, quality<span style="color:#ff79c6">=</span><span style="color:#bd93f9">40</span>, optimize<span style="color:#ff79c6">=</span><span style="color:#ff79c6">True</span>)
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#ff79c6">return</span> compressed_path
</span></span></code></pre></div><h2 id="技术难点与解决方案">技术难点与解决方案</h2>
<h3 id="1-反爬虫对抗">1. 反爬虫对抗</h3>
<p><strong>问题</strong>：淘宝有完善的反爬虫机制</p>
<p><strong>解决方案</strong>：</p>
<ul>
<li>使用DrissionPage框架</li>
<li>设置随机延迟</li>
<li>模拟真实用户行为</li>
</ul>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#282a36;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#6272a4"># 随机延迟模拟人类行为</span>
</span></span><span style="display:flex;"><span><span style="color:#ff79c6">import</span> random
</span></span><span style="display:flex;"><span>time<span style="color:#ff79c6">.</span>sleep(random<span style="color:#ff79c6">.</span>uniform(<span style="color:#bd93f9">1</span>, <span style="color:#bd93f9">3</span>))
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#6272a4"># 滚动页面</span>
</span></span><span style="display:flex;"><span><span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>browser<span style="color:#ff79c6">.</span>scroll_to_bottom()
</span></span></code></pre></div><h3 id="2-图片上传处理">2. 图片上传处理</h3>
<p><strong>问题</strong>：淘宝使用隐藏的file input</p>
<p><strong>解决方案</strong>：</p>
<ul>
<li>使用JavaScript直接设置文件路径</li>
<li>监听上传进度事件</li>
</ul>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#282a36;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"><code class="language-text" data-lang="text"><span style="display:flex;"><span># 直接设置文件路径
</span></span><span style="display:flex;"><span>file_input = self.browser.ele(&#39;css:#image-search-custom-file-input&#39;)
</span></span><span style="display:flex;"><span>self.browser.run_js(f&#34;arguments[0].value = &#39;{image_path}&#39;&#34;, file_input)
</span></span></code></pre></div><h3 id="3-数据解析复杂性">3. 数据解析复杂性</h3>
<p><strong>问题</strong>：淘宝API返回JSONP格式，结构复杂</p>
<p><strong>解决方案</strong>：</p>
<ul>
<li>递归解析JSON结构</li>
<li>使用多种字段别名匹配</li>
<li>建立数据质量评分</li>
</ul>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#282a36;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"><code class="language-python" data-lang="python"><span style="display:flex;"><span><span style="color:#ff79c6">def</span> <span style="color:#50fa7b">find_items_recursively</span>(<span style="font-style:italic">self</span>, obj):
</span></span><span style="display:flex;"><span>    <span style="color:#f1fa8c">&#34;&#34;&#34;递归查找商品数据&#34;&#34;&#34;</span>
</span></span><span style="display:flex;"><span>    <span style="color:#ff79c6">if</span> <span style="color:#8be9fd;font-style:italic">isinstance</span>(obj, <span style="color:#8be9fd;font-style:italic">dict</span>) <span style="color:#ff79c6">and</span> <span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>_is_product_item(obj):
</span></span><span style="display:flex;"><span>        <span style="color:#ff79c6">return</span> [<span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>_extract_product_info(obj)]
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#ff79c6">if</span> <span style="color:#8be9fd;font-style:italic">isinstance</span>(obj, <span style="color:#8be9fd;font-style:italic">list</span>):
</span></span><span style="display:flex;"><span>        <span style="color:#ff79c6">for</span> item <span style="color:#ff79c6">in</span> obj:
</span></span><span style="display:flex;"><span>            result <span style="color:#ff79c6">=</span> <span style="font-style:italic">self</span><span style="color:#ff79c6">.</span>find_items_recursively(item)
</span></span><span style="display:flex;"><span>            <span style="color:#ff79c6">if</span> result:
</span></span><span style="display:flex;"><span>                <span style="color:#ff79c6">return</span> result
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#ff79c6">return</span> []
</span></span></code></pre></div><h2 id="项目成果">项目成果</h2>
<h3 id="功能实现">功能实现</h3>
<ul>
<li>✅ 批量图片搜索（15张图片/批次）</li>
<li>✅ 自动数据提取和解析</li>
<li>✅ 多Sheet Excel报告生成</li>
<li>✅ 邮件自动发送</li>
<li>✅ 数据自动清理</li>
</ul>
<h3 id="性能指标">性能指标</h3>
<ul>
<li><strong>处理速度</strong>：15张图片约3分钟</li>
<li><strong>文件大小</strong>：从326MB压缩到16MB</li>
<li><strong>成功率</strong>：95%以上</li>
<li><strong>稳定性</strong>：支持错误重试</li>
</ul>
<h3 id="用户体验">用户体验</h3>
<ul>
<li><strong>一键运行</strong>：<code>python run.py</code></li>
<li><strong>配置简单</strong>：只需配置邮件信息</li>
<li><strong>日志详细</strong>：完整的执行日志</li>
</ul>
<h2 id="项目结构">项目结构</h2>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#282a36;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"><code class="language-text" data-lang="text"><span style="display:flex;"><span>taobao-search/
</span></span><span style="display:flex;"><span>├── run.py                    # 主启动脚本
</span></span><span style="display:flex;"><span>├── src/                      # 源代码
</span></span><span style="display:flex;"><span>│   ├── automation/           # 自动化模块
</span></span><span style="display:flex;"><span>│   ├── email/               # 邮件服务
</span></span><span style="display:flex;"><span>│   ├── excel/               # Excel处理
</span></span><span style="display:flex;"><span>│   └── workflow/            # 工作流程
</span></span><span style="display:flex;"><span>├── config/                  # 配置文件
</span></span><span style="display:flex;"><span>├── IMG_LIST/                # 图片目录
</span></span><span style="display:flex;"><span>└── data/                    # 数据目录
</span></span></code></pre></div><h2 id="使用方法">使用方法</h2>
<ol>
<li><strong>准备图片</strong>：将图片放入<code>IMG_LIST</code>目录</li>
<li><strong>配置邮件</strong>：编辑<code>config/email_config.json</code></li>
<li><strong>一键运行</strong>：<code>python run.py</code></li>
<li><strong>查看结果</strong>：Excel文件在<code>data/exports/</code>目录</li>
</ol>
<h2 id="技术总结">技术总结</h2>
<h3 id="收获">收获</h3>
<ol>
<li><strong>自动化框架选择</strong>：DrissionPage在反爬虫方面表现优秀</li>
<li><strong>数据拦截技术</strong>：mitmproxy是API数据提取的有效方案</li>
<li><strong>图片处理优化</strong>：合理的压缩策略能显著减小文件大小</li>
<li><strong>工作流程设计</strong>：模块化设计便于维护和扩展</li>
</ol>
<h3 id="价值">价值</h3>
<ul>
<li><strong>效率提升</strong>：从手动操作到全自动化，效率提升10倍</li>
<li><strong>质量保证</strong>：自动化处理减少人为错误</li>
<li><strong>可扩展性</strong>：模块化设计便于功能扩展</li>
</ul>
<h2 id="未来优化">未来优化</h2>
<ol>
<li><strong>支持更多平台</strong>：扩展到京东、拼多多等</li>
<li><strong>增加数据分析</strong>：价格趋势分析、竞品对比</li>
<li><strong>优化用户体验</strong>：Web界面、实时进度显示</li>
<li><strong>增强稳定性</strong>：更完善的错误处理</li>
</ol>
<h2 id="结语">结语</h2>
<p>这个项目从需求分析到最终实现，经历了完整的产品开发周期。通过合理的技术选型、模块化的架构设计和完善的错误处理，最终实现了一个稳定可靠的自动化系统。</p>
<p>最大的挑战是反爬虫对抗和数据解析的复杂性，通过不断调试和优化，最终找到了有效的解决方案。</p>
<p>这个项目不仅解决了实际的业务问题，也让我在自动化测试、数据处理、系统架构等方面有了更深入的理解和实践经验。</p>
]]></content:encoded></item></channel></rss>