Sunday, 11 October 2015

知乎 - DIY 问题重定向

今天我们来玩重定向游戏~

先 logout 知乎网页。 然后浏览 http://www.zhihu.com/question/20381025。

就会看见 5 秒后将会跳转信息:


五秒过后, 跳去 www.zhihu.com/question/19694728?rf=20381025:



五秒过后, 又再跳去 www.zhihu.com/question/22447061?rf=19694728:



还没跳完,五秒过后, 又再跳去 www.zhihu.com/question/20135771?rf=22447061:



五秒过后, 跳去 www.zhihu.com/question/24531839?rf=20135771:



终于跳完 liao~ 是不是很好玩叻。当然你会觉得知乎 desgin 到酱真的是。。。

为什么我要强调 logout ? 因为 login 对知乎的重定向很大影响哦。如果没有 login, 以上所有的重定向都是交给了 js 来处理。 如果有 login,第一个重定向交给 302 处理 (所以你看不见第一页, 而是直接跳到第二页。

来证明一下,login 知乎后, 然后打开  http://www.zhihu.com/question/20381025



留意 URL bar, 是不是很奇怪叻, 直接 skip 掉 20381025 跳第二页。哦, 我还没讲上述 URL 的结构。

http://www.zhihu.com/question/问题_ID?rf=上一个重定向问题_ID

就这样, rf 猜到就是 referer 的简称。

不止这个特别,还会出现从什么问题跳转过来的信息。这就是 login 的特别之处。不但 skip 掉第一个 302, 而且会留意从哪里来。

那么如果想看回重定向之前的那页怎么办 ? 只要来得及按 "从问题 xxxxx  跳转而来" 就能看回之前那页, 并且不再重定向。这种设计,用户怎么可能懂呢 ?

login 的影响 inconsistent 的 behavior 还没完。"从问题 xxxxx  跳转而来" 本身就是 URL 加上 nr=1, 即是 no-redirect 的简称。 如果你没有 login, 你看不见 "从问题 xxxxx  跳转而来" 这行字也就算了,可是如果你手动把 URL 加上 nr=1, 比如 www.zhihu.com/question/20381025?nr=1:


还是会再跳的哦, 也就是说 login 不仅仅影响 302+js 重定向以及 block 掉 rf 信息, 而且也不允许 nr 的功能。

当然我是能想到 302+js 的初衷。302 省一次重定向, 然后又不想用户完全不懂第一页的题目, 所以只 302 静悄悄重定向一次, 接下来就用 "从问题 xxxxx  跳转而来" + js 方法重定向下一页。至于为什么要 login only 就感觉像 bug 多一点。

你可能看不懂我讲什么鬼, 简单一句, 如果是 http://www.zhihu.com/question/问题_ID?rf=上一个重定向问题_ID 链接就加上 &nr=1, 改成
http://www.zhihu.com/question/问题_ID?rf=上一个重定向问题_ID&nr=1 ,如果是 http://www.zhihu.com/question/问题_ID 链接就改成 http://www.zhihu.com/question/问题_ID?nr=1, 并且在 login 的情况下,链接是不会自动重定向的哦。

接下来让我们 DIY 两种重定向规则。

先准备 google chrome 浏览器。目标是自制 chrome extension。

准备两个 file, hello_inject.js 以及 manifest.json 放在同一个 directory/folder:



hello_inject.js 源码:

chrome.webRequest.onBeforeRequest.addListener(
        function(details) {
                if (details.url && /^https?:\/\/www.zhihu.com\/question\/\d+/i.test(details.url)) {
                        var url = details.url;
                        var paramName = "nr";
                        var paramValue = "1";

                        var splitAtAnchor = url.split('#');
                        url = splitAtAnchor[0];
                        var anchor = typeof splitAtAnchor[1] === 'undefined' ? '' : '#' + splitAtAnchor[1];
                        if (url.indexOf(paramName + "=") >= 0) {
                                var prefix = url.substring(0, url.indexOf(paramName));
                                var suffix = url.substring(url.indexOf(paramName));
                                suffix = suffix.substring(suffix.indexOf("=") + 1);
                                suffix = (suffix.indexOf("&") >= 0) ? suffix.substring(suffix.indexOf("&")) : "";
                                url = prefix + paramName + "=" + paramValue + suffix;
                        }
                        else {
                                if (url.indexOf("?") < 0)
                                        url += "?" + paramName + "=" + paramValue;
                                else
                                        url += "&" + paramName + "=" + paramValue;
                        }
 
                        console.log("injected url: " + url + anchor);
                        return {redirectUrl: url + anchor};
                }
        },
        {urls: ["<all_urls>"]},
        ["blocking"]);


manifest.json 源码:

{
  "name": "URL Injector",
  "version": "1.0",
  "description": "Injet url.",
  "background": {
    "scripts": ["hello_inject.js"]
     },
   "manifest_version": 2,
   "permissions": [ "contextMenus", "storage", "webRequest", "webRequestBlocking", "tabs", "http://*/*", "https://*/*" ]
}


然后去 chrome://extensions, 右手边打勾 Developer mode, 然后按 Load unpacked extension...:


选你刚才那个 directory/folder, 按 Open:



源码有 console.log 函数, 所以要看 log 就按 background page 打开 debug 窗口,然后选 Console Tab:



login 情况下,然后浏览 http://www.zhihu.com/question/20381025。



yeah, inject nr=1 成功 :)。现在每一跳问题都会自动加上 nr=1 不再跳转。当然, 必须 login 哟。

如果你不要 print log 就前面加上 // 来 comment out console.log,
//console.log("redirect url: " + url + anchor);

如果更改储存 hello_inject.js 源码文件后, 必须在 chrome://extensions/ 页面, Ctrl+r 更新 take effect。

另一种玩法是选择直接强迫它 302 redirect。原理就是依赖 login + 把 rf 丢掉。因前面研究过,我们已经懂第一页是 302, 特征就是第一页没有 rf 然后接下来有 rf。

hello_inject.js 源码:
chrome.webRequest.onBeforeRequest.addListener(
        function(details) {
                if (details.url && /^https?:\/\/www.zhihu.com\/question\//i.test(details.url)) {
                        var url = details.url;
                        var base = url.substring(0, url.indexOf("?"));
                        var queryString = url.substring(url.indexOf("?") + 1);
                        var params = queryString.split("&");
                        var valid_param = 0;
                        var invalid_param = 0;
                        for (var i = 0; i < params.length; i++) {
                                var p = params[i].split("=");
                                if ((params[i] !== p[0]) && ( p[0] == "rf" )) {
                                        invalid_param+=1;
                                } else {
                                        if (valid_param !== 0) {
                                                base+=params[i];
                                                valid_param+=1;
                                        } else {
                                                base = base + "?" + params[i];
                                        }
                                }
                        }
                        if (invalid_param !== 0) {
                                console.log("redirect base: " + base);
                                return {redirectUrl: base};
                        }
                }
        },
        {urls: ["<all_urls>"]},
        ["blocking"]);


Ctrl+r 更新 extension, login 情况下浏览 http://www.zhihu.com/question/20381025 就会直接 skip 到最后一页 24531839:




如果 Inspect element 检测, 就能看到 302 -> 307(Internal Redirect) -> 302 -> 307 -> 302 -> 307 -> 302 -> 307 skip 到最后一页 200 OK。省下中间的 download。











Sunday, 4 October 2015

pdf - 如何丢掉图片


声明: 看过我的 blog 的人应该懂我几乎都不写这种教你下载看不到 source 或非官方的 exe,我写这篇文章主要是急救个朋友吧了,那个 exe 是不是恶意的我可不负责。

选择 1: 下载安装开源的 libreoffice。缺点是字体有些会走位。

选择 2: 下载安装 Acrobat Pro XI 付费才能用的 edit pdf 功能。本人参考这个教学后,整理出比较小白的步骤。


步骤:
1.  这里下载注册器。点击 Click to Download Content 即可。


2. 下载后,先不要管它。等下才打开 xf-aarpxi。


3. 去这里下载 Acrobat XI Pro。选 File 1 of 1 即可。


4.  浏览 C:\Windows\System32, 找 cmd 这个文件。

5. 滑鼠对着 cmd, 右键选 Run as administrator。


6. [cmd] 输入 cd drivers\etc ,按 Enter 键。再输入 notepad hosts , 按 Enter 键。


7. [notepad] 把下面的内容 copy-paste 去刚才那个 notepad 的最下面。

127.0.0.1 activate.adobe.com
127.0.0.1 activate.adobe.com
127.0.0.1 practivate.adobe.com
127.0.0.1 ereg.adobe.com
127.0.0.1 activate.wip3.adobe.com
127.0.0.1 wip3.adobe.com
127.0.0.1 3dns-3.adobe.com
127.0.0.1 3dns-2.adobe.com
127.0.0.1 adobe-dns.adobe.com
127.0.0.1 adobe-dns-2.adobe.com
127.0.0.1 adobe-dns-3.adobe.com
127.0.0.1 ereg.wip3.adobe.com
127.0.0.1 activate-sea.adobe.com
127.0.0.1 wwis-dubc1-vip60.adobe.com
127.0.0.1 activate-sjc0.adobe.com
127.0.0.1 lmlicenses.wip4.adobe.com
127.0.0.1 lm.licenses.adobe.com


8. [notepad] Save 起来。然后可以关掉 notepad 和 cmd 了。


9. 如果你是用 wifi 就关掉你的 wifi 。 如果是用 network cable 就拔掉 cable。


10.. 双击之前 unrar 的 xf-aarpxi 打开 [注册器] standby。


11. 双击之前下载的 AcrobatPro_11_Web_WWMUI 打开 [acrobat]。 点选 Overwrite。按 Next。


12. [acrobat] Preparing files...Extracting, 等它跑完。


13. [acrobat] Files Are Ready。已经勾上 Launch Adobe Acrobat XI。按 Finish。

14.  如果出现 [UAC dialog] , 按 Yes 允许 program 更改。


15. [Setup] English (United States) -> 按 OK。稍等。


16.  [Setup] InstallShield----, 按 Next。

17. [Setup] 打勾 Make Adobe Acrobat my default PDF viewer。按 Next。


18. [Setup] User Name 填上 email 。点选 "I have a serial number"。

19. [注册器] 去注册器按 "Generate" 左下角的按钮。Serial 栏出现长长的 code。右键 higlight copy 那个 code。


20.  [Setup] paste 那个 code 进去 Serial Number: 栏。按 Next。

21. [Setup] Notcie: Adobe License Activation。按 Next。


22. [Setup] Typical/Complete/Custom。多好过少,点选 Complete。 按 Next。


23. [Setup] Destination Folder。按 Next。


24. [Setup] Ready to Install the Program。按 Install。

25.  [Setup] Status: Copying new files/Patching files/...。 等它跑完。

26.  [Setup] Setup Completed. 按 Finish。


27. 在 desktop 双击安装好的 Adobe Acrobat XI Pro shotcut。


28. [Acrobat XI Pro] Adobe Software License Agreement。 按 Accept。




29.  [Acrobat XI Pro] Serial Number Validation。 请按 "Having trouble connecting to the internet?" 蓝色字。不要按 Validate。


30. [Acrobat XI Pro] No Internet Connection。按 "Offline Activation"。


31. [Acrobat XI Pro] Offline Activation。按 "Generate Request Code"。


32. [Acrobat XI Pro] Offline Activation。滑鼠 highlight/右键 copy "Request Code:" 下面的长长 code。


33. [注册器] 去注册器。弄干净第二行的 "Request:" 栏。右键 paste 进去第二行。


34. [注册器] 按 "Generate"。第三行的 "Activation:" 栏出现长长 code。右键 higlight copy 那个 code。


35. [Acrobat XI Pro]  paste 进去 "Response Code:" 栏。按 Activate。

36. [Acrobat XI Pro] Offline Activation Complete。按 Launch。


37. [Acrobat XI Pro] Help 菜单看得出已经 Activate 了。中间 Select a Task, 选 Edit PDF。选你要 edit 的 pdf 文件。


38. [Acrobat XI Pro] Ctrl+滑鼠轮子放大。 对着图片按会出现格子。按 Delete 键即可丢掉图片。


39. 对着 pdf 右键可以直接打开 Edit。

Friday, 26 June 2015

bash history 时间分组


灵感来源:
不同 patch 不同 project。突然间要重新使用之前 patch to 某一个 project 的 commands, 才发现很死鬼乱水。因为那些 command paths 都是 similar 的, 除非有办法可以快速一眼望去哪些 commands 是坐落在哪一个时间段。

先看一下我的 ~/.bashrc 的一部分:

shopt -s histverify
shopt -s histappend
HISTTIMEFORMAT="%Y/%m/%d %T "
alias histime='history'
alias hisdefault='(HISTTIMEFORMAT=""; history;)'
alias h=hisdefault
alias htime=histime
HISTFILESIZE= 
HISTSIZE= 
HISTCONTROL=ignoreboth



上面的设置现阶段用得还 ok (除了 HISTSIZE=, 看我的解答), 但是 history 增加到成千上万的时候,那些 timestamp 就变得比较 meaningless, 一眼望去都没想过要去看时间。

通常我 type similar 的 command 时, 要 search 的时候只能靠 grep (Tab + Page Up/Down 也可以), 可是那个时间由于 multiple tabs 没有排好好,而且不懂哪里一个打哪里一个, 这个是我跑 histime 的截屏:


所以我就着手写了 hisblock.sh。

我用的 utilities 也是 builtin 的。可是我的 history 有 4 万多条 lines,shell script 一个个去 parse 很不 effcient, 慢到鬼样。

我也试过用很复杂的 binary tree , 除2 再 除 2 循环找到 time range, 但是越做越复杂,很可能有 bug 。

不过后来想想, 直接伪造两条纪录当浮标, 用 append 方式塞进去 history, 就能够直接 sort 了, 然后用 grep 找到浮标的 index, 再用 sed 拉出来, 完全不用烦什么 binary tree sorting 找时间。


然后分两种选择, either 用 default 的 B 或 加上 D。

Block 意思是连续 highlight color 只能在特定的时间范围内。 比如说你 pass 3600 秒, 在 1 小时内的 command lines 都是同一个黄色, 然后下一个小时换成红色, 以此论推。

Distance 则是 highlight color 会检查每一条 command 和下一条 command 的距离, 一旦超过特定时间范围才会变下一种 color。比如说你 pass 30 秒,如果接下去的每一个 commands 都在 30 秒内发生(不是总共哦,而是 "每条的下一条" 重新 check 30 秒), 那么就会全部同一种颜色。

代码如下:

#!/usr/bin/env bash
#Author: <limkokhole@facebook.com>
fname="hisblock.sh"
h_tmp_f="/tmp/hisblock.log"
h_tmp_f2="/tmp/hissorted.log"
p_usage () {
    echo -e "
        BASIC SYNOPSIS:
                source ${fname} SINGLE_QUOTE [from_date] [from_time] [to_date] [to_time] SINGLE_QUOTE interval_in_seconds [B|D]
        Example Usage:
                . ${fname} '2015/05/21' 120 #entire day
                . ${fname} '01:30:00 07:30:00' 120 #default today
                . ${fname} '2015/05/21 01:30:00 07:30:00' 120 #same day
                . ${fname} '2015/05/21 01:30:00 2015/05/22 12:30:00' 120 B #'B' stands for fixed time Block, default
                . ${fname} '2015/05/21 01:30:00 2015/05/22 12:30:00' 120 #120 seconds
                . ${fname} '2015/05/21 01:30:00 2015/05/22 12:30:00' 15 D #'D' for distance between each history line instead of fixed block
"
}

h_swap () {
    if (( "$start_t" > "$end_t" )); then #swap to ignore "to timestamp" and "from timestamp" arg order
        read start_t end_t <<<"$end_t $start_t"
    fi
}

#u must use source OR dot(like how .bashrc do) to run this script bcoz history corrupted even u do `HISTFILE=~/.bash_history` and `set -o history`
if [[ "$(basename -- "$0")" == "$fname" ]]; then
    echo "Don't run $0, instead please use source OR better use . dot" >&2
    p_usage
    exit #can only `return' from a function or sourced script
fi

do_distance=false
if [ "$#" -eq 3 ]; then
    if [[ "$3" == 'D' ]]; then
        do_distance=true
    fi
    t_block="$2"
elif [ "$#" -eq 2 ]; then
 t_block="$2"
else
    p_usage
    return #sourcing don't use exit
fi

d_atom=(`echo ${1}`)
d_len="${#d_atom[@]}"
if (( "$d_len" == 4 )); then
    start_t="$(date -d "${d_atom[0]} ${d_atom[1]}" +%s)" #start timestamp
    end_t="$(date -d "${d_atom[2]} ${d_atom[3]}" +%s)" #end timestamp
elif (( "$d_len" == 3 )); then
    start_t="$(date -d "${d_atom[0]} ${d_atom[1]}" +%s)" #start timestamp
    end_t="$(date -d "${d_atom[0]} ${d_atom[2]}" +%s)" #end timestamp
elif (( "$d_len" == 2 )); then
    today_d="$(date '+%Y/%m/%d')"
    start_t="$(date -d "${today_d} ${d_atom[0]}" +%s)" #start timestamp
    end_t="$(date -d "${today_d} ${d_atom[1]}" +%s)" #end timestamp
elif (( "$d_len" == 1 )); then
    start_t="$(date -d "${d_atom[0]} 00:00:00" +%s)" #start timestamp
    end_t="$(date -d "${d_atom[0]} 23:59:59" +%s)" #end timestamp
else
    p_usage
    return
fi
if [[ "$start_t" =~ ^[0-9]+$ && "$end_t" =~ ^[0-9]+$ && "$t_block" =~ ^[0-9]+$ ]]; then :; else p_usage; return; fi;
h_swap

HISTTIMEFORMAT="%s %Y/%m/%d %T "
next_t="0"

p_red=$(tput setaf 1)
p_green=$(tput setaf 10)
p_yellow=$(tput setaf 11)
p_blue=$(tput setaf 21)
p_orig=$(tput sgr0)
c_arr=($p_red $p_green $p_yellow)
c_arr_len="${#c_arr[@]}"
color_index=0

history >"$h_tmp_f"
printf "%s\n" "START $start_t `date -d @${start_t}`" >>"$h_tmp_f"
printf "%s\n" "END $end_t `date -d @${end_t}`" >>"$h_tmp_f"
sort -k2 -n "$h_tmp_f" > "$h_tmp_f2"
start_index="$(grep -n "^S" "$h_tmp_f2"|cut -f1 -d: )"
end_index="$(grep -n "^E" "$h_tmp_f2" |cut -f1 -d: )"
set -f #noglob
if [ "$do_distance" = false ] ; then #block
    sed -n $(($end_index + 1))'q;'"$start_index","$end_index"p "$h_tmp_f2" |  while read -r line; do
        h_atom=(`echo "${line}"`)
        curr_t="${h_atom[1]}"
        if (( "$curr_t" > "$next_t" )); then
            printf "%s" "${c_arr[ $(($color_index % $c_arr_len)) ]}"
            ((color_index+=1))
            ((next_t="$curr_t"+"$t_block"))
        fi
        h_tail=( "${h_atom[@]:2}" )
        echo "${h_atom[0]} ${h_tail[@]}"
    done
else #distance
    prev_t=0
    sed -n $(($end_index + 1))'q;'"$start_index","$end_index"p "$h_tmp_f2" |  while read -r line; do
        h_atom=(`echo "${line}"`)
        curr_t="${h_atom[1]}"
        if (( $(($curr_t - $prev_t)) > "$t_block" )); then
            printf "%s" "${c_arr[ $(($color_index % $c_arr_len)) ]}"
            ((color_index+=1))
        fi
        prev_t="$curr_t"
        h_tail=( "${h_atom[@]:2}" )
        echo "${h_atom[0]} ${h_tail[@]}"
    done
fi
set +f #reset glob
printf "%s" "${p_orig}"


然后这个是跑了 hisblock.sh 代码的截屏:


 跟之前的青一色相比, 是不是比较清晰了一些叻 :)



Sunday, 10 May 2015

如何在网页上 copy 字

有一些网站会禁止用户用右键 copy 字体。有些更甚的会禁止用滑鼠 highlight。

比如说这个网站, 用 Firefox 打开 (本文不教其它浏览器哟):



有时候会觉得被网站耍的感觉, 想直接右键 Google search 相关内容深入了解, 不能。给一大堆联络地址, URL, 要手抄不成 ?

如果有心人要 copy, 不会天真以为 TA 找不到方法 ? 这种措施非但阻止不了有心人用程序复制网页, 且只会为难一般用户。

通常我们想到的第一个解决方法是 disable javascript, 但是 disable liao 就不给进:



当然, 虽然 Ctrl+S 被禁止, 但是要 Save Page As 还是很容易的:


但是用滑鼠 navigate 不到,选不到字, 而且右键又没 menu 出来。

但是其实呢按 Shift+F10 键是能打开 menu 的哦(这个 shortcut key 是在 destkop, file explore...等等其它地方都能用的),  缺点只是无法选地方打开:



其实更好的打开的右键做法,是输入浏览 about:config 网址。


Enter 选 I'll be careful, I promise!, 就会看到很多个 listing, 在 Search 那里输入 context 关键字(如果没有自动搜索就按 Enter):


在 dom.event.contextmenu.enabled 那边, 双击那个 value 栏目的 true , 就会改成 false



就可以随时随地按右键了。


打开的右键,你可以选择 View Page Source 或 Inspect Element 从那里 Copy 字体。但是麻烦罗, 又要看代码。

所以这里有懒人包做法,直接去选 View ->  Page Style -> No Style


然后就变成可以选了,缺点是 format 走位(其实我觉得更好看,字体深色而且宽)。当然这只是选 highlight, 要使用 Context Menu, 就跟着之前的讲解。


如果要维持 format 呢 ? 这里有一个终极的方法。

打开你平时用的 File Explorer, 然后 Ctrl+L 输入  %appdata%\Mozilla\Firefox


按 Enter, 就会来到这里:


然后选 Profiles -> 随机字.default folder, 来到:


你电脑还没有 chrome 这个 folder, 所以你要右键 New- > Folder 新建 chrome

进去那个 chrome folder, 里面是空的, 要新建一个 css 文件。

So, Win+R 键输入打开 notepad:


输入以下内容:

* { -moz-user-select: text !important;
* user-select: text !important; }

如图所示:

然后选 File -> Save As... 在刚才那个 Chrome 的 file path。

如果你背不上来那个 Chrome 的 path, 你可以 Ctrl+L 然后 Ctrl+C copy 那个 path, 然后在这里 Ctrl+L 然后 Ctrl+P paste 进来, Enter, 就能跳到那个 filepath 了。



这里 file name 要放 userContent.css , 然后 Save as type 选 All Files。

然后就搞掂 liao, 关掉所有 Firefox 然后开过。浏览这个网页, 用滑鼠 highlight 就变成通畅无阻了 :)



小白 Bonus:
除了用右键 copy, 你也可以highlight, 然后用滑鼠中间轮子 click 一下就 copy 了,要 paste 时就用滑鼠中间轮子 click 一下即可。至于 ctrl+c 被禁止我暂时想不到方便的解决方法。我不是 web developer 叻 :)

[2018 更新]:
以上的方法在佳礼中文网不能 works, 不过你可以制作自己的火狐扩展,  如:

xb@dnxb:~/firefox-enable-selection$ cat manifest.json
{
  "manifest_version": 2,
  "name": "firefox-enable-selection",
  "version": "1.0",

  "content_scripts": [
    {
      "matches": ["https://*cari.com.my/*"],
      "js": ["selection.js"]
    }
  ]
}


xb@dnxb:~/firefox-enable-selection$ cat selection.js
Array.from(document.querySelectorAll('div[onmousedown]')).forEach((element,index) =>
{
     element.onmousedown=function(){return true}
});
Array.from(document.querySelectorAll('div[onselectstart]')).forEach((element,index) =>
{
     element.onselectstart=function(){return true}


});
xb@dnxb:~/Downloads/misc/firefox-enable-selection$

因为佳礼的某些 div class 有 `<div class="d" onmousedown="return false;" onselectstart="return false;">` ,
此方法就是循环替代该 elemenets。

或者用我这个扩展。