首頁>技術>

背景

公司有個渠道系統,專門對接三方渠道使用,沒有什麼業務邏輯,主要是轉換報文和引數校驗之類的工作,起著一個承上啟下的作用。

最近在最佳化介面的響應時間,優化了程式碼之後,但是時間還是達不到要求;有一個詭異的 100ms 左右的耗時問題,在介面中列印了請求處理時間後,和呼叫方的響應時間還有差了 100ms 左右。比如程式裡記錄 150ms,但是呼叫方等待時間卻為 250ms 左右。

下面記錄下當時詳細的定位 & 解決流程(其實解決很簡單,關鍵在於怎麼定位並找到解決問題的方法)。

定位過程1. 分析程式碼

渠道系統是一個常見的 Spring-boot web 工程,使用了整合的 tomcat。分析了程式碼之後,發現並沒有特殊的地方,沒有特殊的過濾器或者攔截器,所以初步排除是業務程式碼問題。

2. 分析呼叫流程

出現這個問題之後,首先確認了下介面的呼叫流程。由於是內部測試,所以呼叫流程較少。

Nginx -反向代理-> 渠道系統

公司是雲伺服器,網路走的也是雲的內網。由於不明確問題的原因,所以用排除法,首先確認伺服器網路是否有問題。

先確認傳送端到 Nginx Host 是否有問題:

[jboss@VM_0_139_centos ~]$ ping 10.0.0.139PING 10.0.0.139 (10.0.0.139) 56(84) bytes of data.64 bytes from 10.0.0.139: icmp_seq=1 ttl=64 time=0.029 ms64 bytes from 10.0.0.139: icmp_seq=2 ttl=64 time=0.041 ms64 bytes from 10.0.0.139: icmp_seq=3 ttl=64 time=0.040 ms64 bytes from 10.0.0.139: icmp_seq=4 ttl=64 time=0.040 ms

從 ping 結果上看,傳送端到 Nginx 主機的延遲是無問題的,接下來檢視 Nginx 到渠道系統的網路。

# 由於日誌是沒問題的,這裡直接複製上面日誌了[jboss@VM_0_139_centos ~]$ ping 10.0.0.139PING 10.0.0.139 (10.0.0.139) 56(84) bytes of data.64 bytes from 10.0.0.139: icmp_seq=1 ttl=64 time=0.029 ms64 bytes from 10.0.0.139: icmp_seq=2 ttl=64 time=0.041 ms64 bytes from 10.0.0.139: icmp_seq=3 ttl=64 time=0.040 ms64 bytes from 10.0.0.139: icmp_seq=4 ttl=64 time=0.040 ms

從 ping 結果上看,Nginx 到渠道系統伺服器網路延遲也是沒問題的。

既然網路看似沒問題,那麼可以繼續排除法,砍掉 Nginx,客戶端直接再渠道系統的伺服器上,透過迴環地址(localhost)直連,避免經過網絡卡/dns,縮小問題範圍看看能否復現(這個應用和地址是我後期模擬的,測試的是一個空介面):

[jboss@VM_10_91_centos tmp]$ curl -w "@curl-time.txt" http://127.0.0.1:7744/sendsuccess              http: 200               dns: 0.001s          redirect: 0.000s      time_connect: 0.001s   time_appconnect: 0.000s  time_pretransfer: 0.001stime_starttransfer: 0.073s     size_download: 7bytes    speed_download: 95.000B/s                  ----------        time_total: 0.073s 請求總耗時

從 curl 日誌上看,透過迴環地址呼叫一個空介面耗時也有 73ms。這就奇怪了,跳過了中間所有呼叫節點(包括過濾器 & 攔截器之類),直接請求應用一個空介面,都有 73ms 的耗時,再請求一次看看:

[jboss@VM_10_91_centos tmp]$ curl -w "@curl-time.txt" http://127.0.0.1:7744/sendsuccess              http: 200               dns: 0.001s          redirect: 0.000s      time_connect: 0.001s   time_appconnect: 0.000s  time_pretransfer: 0.001stime_starttransfer: 0.003s     size_download: 7bytes    speed_download: 2611.000B/s                  ----------        time_total: 0.003s

更奇怪的是,第二次請求耗時就正常了,變成了 3ms。經查閱資料,linux curl 是預設開啟 http keep-alive的(Keep-Alive 的介紹可以參考我的另一篇文章)。就算不開啟 keep-alive,每次重新 handshake,也不至於需要 70ms。

經過不斷分析測試發現,連續請求的話時間就會很短,每次請求只需要幾毫秒,但是如果隔一段時間再請求,就會花費 70ms 以上。

從這個現象猜想,可能是某些快取機制導致的,連續請求因為有快取,所以速度快,時間長快取失效後導致時間長。

那麼這個問題點到底在哪一層呢?tomcat 層還是 spring-webmvc 呢?

光猜想定位不了問題,還是得實際測試一下,把渠道系統的程式碼放到本地 IDE 裡啟動測試能否復現。

但是匯入本地 IDE 後,在 IDE 中啟動後並不能復現問題,並沒有 70+ms 的延遲問題。這下頭疼了,本地無法復現,不能 Debug,由於問題點不在業務程式碼,也不能透過加日誌的方式來 Debug。

這時候可以祭出神器 Arthas 了

3. Arthas 分析問題

Arthas 是 Alibaba 開源的 Java 診斷工具,深受開發者喜愛。當你遇到以下類似問題而束手無策時,Arthas 可以幫助你解決:

這個類從哪個 jar 包載入的?為什麼會報各種類相關的 Exception?我改的程式碼為什麼沒有執行到?難道是我沒 commit?分支搞錯了?遇到問題無法在線上 debug,難道只能透過加日誌再重新發布嗎?線上遇到某個使用者的資料處理有問題,但線上同樣無法 debug,線下無法重現!是否有一個全域性視角來檢視系統的執行狀況?有什麼辦法可以監控到 JVM 的實時執行狀態?······

上面是 Arthas 的官方簡介,這次我只需要用他的一個小功能 trace。動態計算方法呼叫路徑和時間,這樣就可以定位時間在哪個地方被消耗了。

trace 方法內部呼叫路徑,並輸出方法路徑上的每個節點上耗時。

trace 命令能主動搜尋 class-pattern/method-pattern 。

對應的方法呼叫路徑,渲染和統計整個呼叫鏈路上的所有效能開銷和追蹤呼叫鏈路。

有了神器,那麼該追蹤什麼方法呢?由於我對 Tomcat 原始碼不是很熟,所以只能從 spring mvc 下手,先來 trace 一下 spring mvc 的入口:

[arthas@24851]$ trace org.springframework.web.servlet.DispatcherServlet *Press Q or Ctrl+C to abort.Affect(class-cnt:1 , method-cnt:44) cost in 508 ms.`---ts=2019-09-14 21:07:44;thread_name=http-nio-7744-exec-2;id=11;is_daemon=true;priority=5;TCCL=org.springframework.boot.web.embedded.tomcat.TomcatEmbeddedWebappClassLoader@7c136917    `---[2.952142ms] org.springframework.web.servlet.DispatcherServlet:buildLocaleContext()`---ts=2019-09-14 21:07:44;thread_name=http-nio-7744-exec-2;id=11;is_daemon=true;priority=5;TCCL=org.springframework.boot.web.embedded.tomcat.TomcatEmbeddedWebappClassLoader@7c136917    `---[18.08903ms] org.springframework.web.servlet.DispatcherServlet:doService()        +---[0.041346ms] org.apache.commons.logging.Log:isDebugEnabled() #889        +---[0.022398ms] org.springframework.web.util.WebUtils:isIncludeRequest() #898        +---[0.014904ms] org.springframework.web.servlet.DispatcherServlet:getWebApplicationContext() #910        +---[1.071879ms] javax.servlet.http.HttpServletRequest:setAttribute() #910        +---[0.020977ms] javax.servlet.http.HttpServletRequest:setAttribute() #911        +---[0.017073ms] javax.servlet.http.HttpServletRequest:setAttribute() #912        +---[0.218277ms] org.springframework.web.servlet.DispatcherServlet:getThemeSource() #913        |   `---[0.137568ms] org.springframework.web.servlet.DispatcherServlet:getThemeSource()        |       `---[min=0.00783ms,max=0.014251ms,total=0.022081ms,count=2] org.springframework.web.servlet.DispatcherServlet:getWebApplicationContext() #782        +---[0.019363ms] javax.servlet.http.HttpServletRequest:setAttribute() #913        +---[0.070694ms] org.springframework.web.servlet.FlashMapManager:retrieveAndUpdate() #916        +---[0.01839ms] org.springframework.web.servlet.FlashMap:<init>() #920        +---[0.016943ms] javax.servlet.http.HttpServletRequest:setAttribute() #920        +---[0.015268ms] javax.servlet.http.HttpServletRequest:setAttribute() #921        +---[15.050124ms] org.springframework.web.servlet.DispatcherServlet:doDispatch() #925        |   `---[14.943477ms] org.springframework.web.servlet.DispatcherServlet:doDispatch()        |       +---[0.019135ms] org.springframework.web.context.request.async.WebAsyncUtils:getAsyncManager() #953        |       +---[2.108373ms] org.springframework.web.servlet.DispatcherServlet:checkMultipart() #960        |       |   `---[2.004436ms] org.springframework.web.servlet.DispatcherServlet:checkMultipart()        |       |       `---[1.890845ms] org.springframework.web.multipart.MultipartResolver:isMultipart() #1117        |       +---[2.054361ms] org.springframework.web.servlet.DispatcherServlet:getHandler() #964        |       |   `---[1.961963ms] org.springframework.web.servlet.DispatcherServlet:getHandler()        |       |       +---[0.02051ms] java.util.List:iterator() #1183        |       |       +---[min=0.003805ms,max=0.009641ms,total=0.013446ms,count=2] java.util.Iterator:hasNext() #1183        |       |       +---[min=0.003181ms,max=0.009751ms,total=0.012932ms,count=2] java.util.Iterator:next() #1183        |       |       +---[min=0.005841ms,max=0.015308ms,total=0.021149ms,count=2] org.apache.commons.logging.Log:isTraceEnabled() #1184        |       |       `---[min=0.474739ms,max=1.19145ms,total=1.666189ms,count=2] org.springframework.web.servlet.HandlerMapping:getHandler() #1188        |       +---[0.013071ms] org.springframework.web.servlet.HandlerExecutionChain:getHandler() #971        |       +---[0.372236ms] org.springframework.web.servlet.DispatcherServlet:getHandlerAdapter() #971        |       |   `---[0.280073ms] org.springframework.web.servlet.DispatcherServlet:getHandlerAdapter()        |       |       +---[0.004804ms] java.util.List:iterator() #1224        |       |       +---[0.003668ms] java.util.Iterator:hasNext() #1224        |       |       +---[0.003038ms] java.util.Iterator:next() #1224        |       |       +---[0.006451ms] org.apache.commons.logging.Log:isTraceEnabled() #1225        |       |       `---[0.012683ms] org.springframework.web.servlet.HandlerAdapter:supports() #1228        |       +---[0.012848ms] javax.servlet.http.HttpServletRequest:getMethod() #974        |       +---[0.013132ms] java.lang.String:equals() #975        |       +---[0.003025ms] org.springframework.web.servlet.HandlerExecutionChain:getHandler() #977        |       +---[0.008095ms] org.springframework.web.servlet.HandlerAdapter:getLastModified() #977        |       +---[0.006596ms] org.apache.commons.logging.Log:isDebugEnabled() #978        |       +---[0.018024ms] org.springframework.web.context.request.ServletWebRequest:<init>() #981        |       +---[0.017869ms] org.springframework.web.context.request.ServletWebRequest:checkNotModified() #981        |       +---[0.038542ms] org.springframework.web.servlet.HandlerExecutionChain:applyPreHandle() #986        |       +---[0.00431ms] org.springframework.web.servlet.HandlerExecutionChain:getHandler() #991        |       +---[4.248493ms] org.springframework.web.servlet.HandlerAdapter:handle() #991        |       +---[0.014805ms] org.springframework.web.context.request.async.WebAsyncManager:isConcurrentHandlingStarted() #993        |       +---[1.444994ms] org.springframework.web.servlet.DispatcherServlet:applyDefaultViewName() #997        |       |   `---[0.067631ms] org.springframework.web.servlet.DispatcherServlet:applyDefaultViewName()        |       +---[0.012027ms] org.springframework.web.servlet.HandlerExecutionChain:applyPostHandle() #998        |       +---[0.373997ms] org.springframework.web.servlet.DispatcherServlet:processDispatchResult() #1008        |       |   `---[0.197004ms] org.springframework.web.servlet.DispatcherServlet:processDispatchResult()        |       |       +---[0.007074ms] org.apache.commons.logging.Log:isDebugEnabled() #1075        |       |       +---[0.005467ms] org.springframework.web.context.request.async.WebAsyncUtils:getAsyncManager() #1081        |       |       +---[0.004054ms] org.springframework.web.context.request.async.WebAsyncManager:isConcurrentHandlingStarted() #1081        |       |       `---[0.011988ms] org.springframework.web.servlet.HandlerExecutionChain:triggerAfterCompletion() #1087        |       `---[0.004015ms] org.springframework.web.context.request.async.WebAsyncManager:isConcurrentHandlingStarted() #1018        +---[0.005055ms] org.springframework.web.context.request.async.WebAsyncUtils:getAsyncManager() #928        `---[0.003422ms] org.springframework.web.context.request.async.WebAsyncManager:isConcurrentHandlingStarted() #928
[jboss@VM_10_91_centos tmp]$ curl -w "@curl-time.txt" http://127.0.0.1:7744/sendsuccess              http: 200               dns: 0.001s          redirect: 0.000s      time_connect: 0.001s   time_appconnect: 0.000s  time_pretransfer: 0.001stime_starttransfer: 0.115s     size_download: 7bytes    speed_download: 60.000B/s                  ----------        time_total: 0.115s

本次呼叫,呼叫端時間花費 115 ms,但是從 arthas trace 上看,spring mvc 只消耗了 18ms,那麼剩下的 97ms 去哪了呢?

本地測試後已經可以排除 spring mvc 的問題了,最後也是唯一可能出問題的點就是 tomcat。

可是本人並不熟悉 tomcat 中的原始碼,就連請求入口都不清楚,tomcat 裡需要 trace 的類都不好找。。。

不過沒關係,有神器 Arthas,可以透過 stack 命令來反向查詢呼叫路徑,以org.springframework.web.servlet.DispatcherServlet 作為引數:

stack 輸出當前方法被呼叫的呼叫路徑。

很多時候我們都知道一個方法被執行,但這個方法被執行的路徑非常多,或者你根本就不知道這個方法是從那裡被執行了,此時你需要的是 stack 命令。

[arthas@24851]$ stack org.springframework.web.servlet.DispatcherServlet *Press Q or Ctrl+C to abort.Affect(class-cnt:1 , method-cnt:44) cost in 495 ms.ts=2019-09-14 21:15:19;thread_name=http-nio-7744-exec-5;id=14;is_daemon=true;priority=5;TCCL=org.springframework.boot.web.embedded.tomcat.TomcatEmbeddedWebappClassLoader@7c136917    @org.springframework.web.servlet.FrameworkServlet.processRequest()        at org.springframework.web.servlet.FrameworkServlet.doGet(FrameworkServlet.java:866)        at javax.servlet.http.HttpServlet.service(HttpServlet.java:635)        at org.springframework.web.servlet.FrameworkServlet.service(FrameworkServlet.java:851)        at javax.servlet.http.HttpServlet.service(HttpServlet.java:742)        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:231)        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)        at org.apache.tomcat.websocket.server.WsFilter.doFilter(WsFilter.java:52)        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)        at org.springframework.web.filter.RequestContextFilter.doFilterInternal(RequestContextFilter.java:99)        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)        at org.springframework.web.filter.HttpPutFormContentFilter.doFilterInternal(HttpPutFormContentFilter.java:109)        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)        at org.springframework.web.filter.HiddenHttpMethodFilter.doFilterInternal(HiddenHttpMethodFilter.java:81)        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)        at org.springframework.web.filter.CharacterEncodingFilter.doFilterInternal(CharacterEncodingFilter.java:200)        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)        at org.apache.catalina.core.StandardWrapperValve.invoke(StandardWrapperValve.java:198)        at org.apache.catalina.core.StandardContextValve.invoke(StandardContextValve.java:96)        at org.apache.catalina.authenticator.AuthenticatorBase.invoke(AuthenticatorBase.java:496)        at org.apache.catalina.core.StandardHostValve.invoke(StandardHostValve.java:140)        at org.apache.catalina.valves.ErrorReportValve.invoke(ErrorReportValve.java:81)        at org.apache.catalina.core.StandardEngineValve.invoke(StandardEngineValve.java:87)        at org.apache.catalina.connector.CoyoteAdapter.service(CoyoteAdapter.java:342)        at org.apache.coyote.http11.Http11Processor.service(Http11Processor.java:803)        at org.apache.coyote.AbstractProcessorLight.process(AbstractProcessorLight.java:66)        at org.apache.coyote.AbstractProtocol$ConnectionHandler.process(AbstractProtocol.java:790)        at org.apache.tomcat.util.net.NioEndpoint$SocketProcessor.doRun(NioEndpoint.java:1468)        at org.apache.tomcat.util.net.SocketProcessorBase.run(SocketProcessorBase.java:49)        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)        at org.apache.tomcat.util.threads.TaskThread$WrappingRunnable.run(TaskThread.java:61)        at java.lang.Thread.run(Thread.java:748)ts=2019-09-14 21:15:19;thread_name=http-nio-7744-exec-5;id=14;is_daemon=true;priority=5;TCCL=org.springframework.boot.web.embedded.tomcat.TomcatEmbeddedWebappClassLoader@7c136917    @org.springframework.web.servlet.DispatcherServlet.doService()        at org.springframework.web.servlet.FrameworkServlet.processRequest(FrameworkServlet.java:974)        at org.springframework.web.servlet.FrameworkServlet.doGet(FrameworkServlet.java:866)        at javax.servlet.http.HttpServlet.service(HttpServlet.java:635)        at org.springframework.web.servlet.FrameworkServlet.service(FrameworkServlet.java:851)        at javax.servlet.http.HttpServlet.service(HttpServlet.java:742)        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:231)        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)        at org.apache.tomcat.websocket.server.WsFilter.doFilter(WsFilter.java:52)        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)        at org.springframework.web.filter.RequestContextFilter.doFilterInternal(RequestContextFilter.java:99)        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)        at org.springframework.web.filter.HttpPutFormContentFilter.doFilterInternal(HttpPutFormContentFilter.java:109)        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)        at org.springframework.web.filter.HiddenHttpMethodFilter.doFilterInternal(HiddenHttpMethodFilter.java:81)        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)        at org.springframework.web.filter.CharacterEncodingFilter.doFilterInternal(CharacterEncodingFilter.java:200)        at org.springframework.web.filter.OncePerRequestFilter.doFilter(OncePerRequestFilter.java:107)        at org.apache.catalina.core.ApplicationFilterChain.internalDoFilter(ApplicationFilterChain.java:193)        at org.apache.catalina.core.ApplicationFilterChain.doFilter(ApplicationFilterChain.java:166)        at org.apache.catalina.core.StandardWrapperValve.invoke(StandardWrapperValve.java:198)        at org.apache.catalina.core.StandardContextValve.invoke(StandardContextValve.java:96)        at org.apache.catalina.authenticator.AuthenticatorBase.invoke(AuthenticatorBase.java:496)        at org.apache.catalina.core.StandardHostValve.invoke(StandardHostValve.java:140)        at org.apache.catalina.valves.ErrorReportValve.invoke(ErrorReportValve.java:81)        at org.apache.catalina.core.StandardEngineValve.invoke(StandardEngineValve.java:87)        at org.apache.catalina.connector.CoyoteAdapter.service(CoyoteAdapter.java:342)        at org.apache.coyote.http11.Http11Processor.service(Http11Processor.java:803)        at org.apache.coyote.AbstractProcessorLight.process(AbstractProcessorLight.java:66)        at org.apache.coyote.AbstractProtocol$ConnectionHandler.process(AbstractProtocol.java:790)        at org.apache.tomcat.util.net.NioEndpoint$SocketProcessor.doRun(NioEndpoint.java:1468)        at org.apache.tomcat.util.net.SocketProcessorBase.run(SocketProcessorBase.java:49)        at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)        at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)        at org.apache.tomcat.util.threads.TaskThread$WrappingRunnable.run(TaskThread.java:61)        at java.lang.Thread.run(Thread.java:748)

從 stack 日誌上可以很直觀的看出 DispatchServlet 的呼叫棧,那麼這麼長的路徑,該 trace 哪個類呢(這裡跳過 spring mvc 中的過濾器的 trace 過程,實際排查的時候也 trace 了一遍,但這詭異的時間消耗不是由這裡過濾器產生的)?有一定經驗的老司機從名字上大概也能猜出來從哪裡下手比較好,那就是org.apache.coyote.http11.Http11Processor.service,從名字上看,http1.1 處理器,這可能是一個比較好的切入點。下面來 trace 一下:

[arthas@24851]$ trace org.apache.coyote.http11.Http11Processor servicePress Q or Ctrl+C to abort.Affect(class-cnt:1 , method-cnt:1) cost in 269 ms.`---ts=2019-09-14 21:22:51;thread_name=http-nio-7744-exec-8;id=17;is_daemon=true;priority=5;TCCL=org.springframework.boot.loader.LaunchedURLClassLoader@20ad9418    `---[131.650285ms] org.apache.coyote.http11.Http11Processor:service()        +---[0.036851ms] org.apache.coyote.Request:getRequestProcessor() #667        +---[0.009986ms] org.apache.coyote.RequestInfo:setStage() #668        +---[0.008928ms] org.apache.coyote.http11.Http11Processor:setSocketWrapper() #671        +---[0.013236ms] org.apache.coyote.http11.Http11InputBuffer:init() #672        +---[0.00981ms] org.apache.coyote.http11.Http11OutputBuffer:init() #673        +---[min=0.00213ms,max=0.007317ms,total=0.009447ms,count=2] org.apache.coyote.http11.Http11Processor:getErrorState() #683        +---[min=0.002098ms,max=0.008888ms,total=0.010986ms,count=2] org.apache.coyote.ErrorState:isError() #683        +---[min=0.002448ms,max=0.007149ms,total=0.009597ms,count=2] org.apache.coyote.http11.Http11Processor:isAsync() #683        +---[min=0.002399ms,max=0.00852ms,total=0.010919ms,count=2] org.apache.tomcat.util.net.AbstractEndpoint:isPaused() #683        +---[min=0.033587ms,max=0.11832ms,total=0.151907ms,count=2] org.apache.coyote.http11.Http11InputBuffer:parseRequestLine() #687        +---[0.005384ms] org.apache.tomcat.util.net.AbstractEndpoint:isPaused() #695        +---[0.007924ms] org.apache.coyote.Request:getMimeHeaders() #702        +---[0.006744ms] org.apache.tomcat.util.net.AbstractEndpoint:getMaxHeaderCount() #702        +---[0.012574ms] org.apache.tomcat.util.http.MimeHeaders:setLimit() #702        +---[0.14319ms] org.apache.coyote.http11.Http11InputBuffer:parseHeaders() #703        +---[0.003997ms] org.apache.coyote.Request:getMimeHeaders() #743        +---[0.026561ms] org.apache.tomcat.util.http.MimeHeaders:values() #743        +---[min=0.002869ms,max=0.01203ms,total=0.014899ms,count=2] java.util.Enumeration:hasMoreElements() #745        +---[0.070114ms] java.util.Enumeration:nextElement() #746        +---[0.010921ms] java.lang.String:toLowerCase() #746        +---[0.008453ms] java.lang.String:contains() #746        +---[0.002698ms] org.apache.coyote.http11.Http11Processor:getErrorState() #775        +---[0.00307ms] org.apache.coyote.ErrorState:isError() #775        +---[0.002708ms] org.apache.coyote.RequestInfo:setStage() #777        +---[0.171139ms] org.apache.coyote.http11.Http11Processor:prepareRequest() #779        +---[0.009349ms] org.apache.tomcat.util.net.SocketWrapperBase:decrementKeepAlive() #794        +---[0.002574ms] org.apache.coyote.http11.Http11Processor:getErrorState() #800        +---[0.002696ms] org.apache.coyote.ErrorState:isError() #800        +---[0.002499ms] org.apache.coyote.RequestInfo:setStage() #802        +---[0.005641ms] org.apache.coyote.http11.Http11Processor:getAdapter() #803        +---[129.868916ms] org.apache.coyote.Adapter:service() #803        +---[0.003859ms] org.apache.coyote.http11.Http11Processor:getErrorState() #809        +---[0.002365ms] org.apache.coyote.ErrorState:isError() #809        +---[0.003844ms] org.apache.coyote.http11.Http11Processor:isAsync() #809        +---[0.002382ms] org.apache.coyote.Response:getStatus() #809        +---[0.002476ms] org.apache.coyote.http11.Http11Processor:statusDropsConnection() #809        +---[0.002284ms] org.apache.coyote.RequestInfo:setStage() #838        +---[0.00222ms] org.apache.coyote.http11.Http11Processor:isAsync() #839        +---[0.037873ms] org.apache.coyote.http11.Http11Processor:endRequest() #843        +---[0.002188ms] org.apache.coyote.RequestInfo:setStage() #845        +---[0.002112ms] org.apache.coyote.http11.Http11Processor:getErrorState() #849        +---[0.002063ms] org.apache.coyote.ErrorState:isError() #849        +---[0.002504ms] org.apache.coyote.http11.Http11Processor:isAsync() #853        +---[0.009808ms] org.apache.coyote.Request:updateCounters() #854        +---[0.002008ms] org.apache.coyote.http11.Http11Processor:getErrorState() #855        +---[0.002192ms] org.apache.coyote.ErrorState:isIoAllowed() #855        +---[0.01968ms] org.apache.coyote.http11.Http11InputBuffer:nextRequest() #856        +---[0.010065ms] org.apache.coyote.http11.Http11OutputBuffer:nextRequest() #857        +---[0.002576ms] org.apache.coyote.RequestInfo:setStage() #870        +---[0.016599ms] org.apache.coyote.http11.Http11Processor:processSendfile() #872        +---[0.008182ms] org.apache.coyote.http11.Http11InputBuffer:getParsingRequestLinePhase() #688        +---[0.0075ms] org.apache.coyote.http11.Http11Processor:handleIncompleteRequestLineRead() #690        +---[0.001979ms] org.apache.coyote.RequestInfo:setStage() #875        +---[0.001981ms] org.apache.coyote.http11.Http11Processor:getErrorState() #877        +---[0.001934ms] org.apache.coyote.ErrorState:isError() #877        +---[0.001995ms] org.apache.tomcat.util.net.AbstractEndpoint:isPaused() #877        +---[0.002403ms] org.apache.coyote.http11.Http11Processor:isAsync() #879        `---[0.006176ms] org.apache.coyote.http11.Http11Processor:isUpgrade() #881

日誌裡有一個 129ms 的耗時點(時間比沒開 arthas 的時候更長是因為 arthas 本身帶來的效能消耗,所以生產環境小心使用),這個就是要找的問題點。

打問題點找到了,那怎麼定位是什麼導致的問題呢,又如何解決呢?

繼續 trace 吧,細化到具體的程式碼塊或者內容。trace 由於效能考慮,不會展示所有的呼叫路徑,如果呼叫路徑過深,只有手動深入 trace,原則就是 trace 耗時長的那個方法:

[arthas@24851]$ trace org.apache.coyote.Adapter servicePress Q or Ctrl+C to abort.Affect(class-cnt:1 , method-cnt:1) cost in 608 ms.`---ts=2019-09-14 21:34:33;thread_name=http-nio-7744-exec-1;id=10;is_daemon=true;priority=5;TCCL=org.springframework.boot.loader.LaunchedURLClassLoader@20ad9418    `---[81.70999ms] org.apache.catalina.connector.CoyoteAdapter:service()        +---[0.032546ms] org.apache.coyote.Request:getNote() #302        +---[0.007148ms] org.apache.coyote.Response:getNote() #303        +---[0.007475ms] org.apache.catalina.connector.Connector:getXpoweredBy() #324        +---[0.00447ms] org.apache.coyote.Request:getRequestProcessor() #331        +---[0.007902ms] java.lang.ThreadLocal:get() #331        +---[0.006522ms] org.apache.coyote.RequestInfo:setWorkerThreadName() #331        +---[73.793798ms] org.apache.catalina.connector.CoyoteAdapter:postParseRequest() #336        +---[0.001536ms] org.apache.catalina.connector.Connector:getService() #339        +---[0.004469ms] org.apache.catalina.Service:getContainer() #339        +---[0.007074ms] org.apache.catalina.Engine:getPipeline() #339        +---[0.004334ms] org.apache.catalina.Pipeline:isAsyncSupported() #339        +---[0.002466ms] org.apache.catalina.connector.Request:setAsyncSupported() #339        +---[6.01E-4ms] org.apache.catalina.connector.Connector:getService() #342        +---[0.001859ms] org.apache.catalina.Service:getContainer() #342        +---[9.65E-4ms] org.apache.catalina.Engine:getPipeline() #342        +---[0.005231ms] org.apache.catalina.Pipeline:getFirst() #342        +---[7.239154ms] org.apache.catalina.Valve:invoke() #342        +---[0.006904ms] org.apache.catalina.connector.Request:isAsync() #345        +---[0.00509ms] org.apache.catalina.connector.Request:finishRequest() #372        +---[0.051461ms] org.apache.catalina.connector.Response:finishResponse() #373        +---[0.007244ms] java.util.concurrent.atomic.AtomicBoolean:<init>() #379        +---[0.007314ms] org.apache.coyote.Response:action() #380        +---[0.004518ms] org.apache.catalina.connector.Request:isAsyncCompleting() #382        +---[0.001072ms] org.apache.catalina.connector.Request:getContext() #394        +---[0.007166ms] java.lang.System:currentTimeMillis() #401        +---[0.004367ms] org.apache.coyote.Request:getStartTime() #401        +---[0.011483ms] org.apache.catalina.Context:logAccess() #401        +---[0.0014ms] org.apache.coyote.Request:getRequestProcessor() #406        +---[min=8.0E-4ms,max=9.22E-4ms,total=0.001722ms,count=2] java.lang.Integer:<init>() #406        +---[0.001082ms] java.lang.reflect.Method:invoke() #406        +---[0.001851ms] org.apache.coyote.RequestInfo:setWorkerThreadName() #406        +---[0.035805ms] org.apache.catalina.connector.Request:recycle() #410        `---[0.007849ms] org.apache.catalina.connector.Response:recycle() #411

一段無聊的手動深入 trace 之後………………

[arthas@24851]$ trace org.apache.catalina.webresources.AbstractArchiveResourceSet getArchiveEntriesPress Q or Ctrl+C to abort.Affect(class-cnt:4 , method-cnt:2) cost in 150 ms.`---ts=2019-09-14 21:36:26;thread_name=http-nio-7744-exec-3;id=12;is_daemon=true;priority=5;TCCL=org.springframework.boot.loader.LaunchedURLClassLoader@20ad9418    `---[75.743681ms] org.apache.catalina.webresources.JarWarResourceSet:getArchiveEntries()        +---[0.025731ms] java.util.HashMap:<init>() #106        +---[0.097729ms] org.apache.catalina.webresources.JarWarResourceSet:openJarFile() #109        +---[0.091037ms] java.util.jar.JarFile:getJarEntry() #110        +---[0.096325ms] java.util.jar.JarFile:getInputStream() #111        +---[0.451916ms] org.apache.catalina.webresources.TomcatJarInputStream:<init>() #113        +---[min=0.001175ms,max=0.001176ms,total=0.002351ms,count=2] java.lang.Integer:<init>() #114        +---[0.00104ms] java.lang.reflect.Method:invoke() #114        +---[0.045105ms] org.apache.catalina.webresources.TomcatJarInputStream:getNextJarEntry() #114        +---[min=5.02E-4ms,max=0.008531ms,total=0.028864ms,count=31] java.util.jar.JarEntry:getName() #116        +---[min=5.39E-4ms,max=0.022805ms,total=0.054647ms,count=31] java.util.HashMap:put() #116        +---[min=0.004452ms,max=34.479307ms,total=74.206249ms,count=31] org.apache.catalina.webresources.TomcatJarInputStream:getNextJarEntry() #117        +---[0.018358ms] org.apache.catalina.webresources.TomcatJarInputStream:getManifest() #119        +---[0.006429ms] org.apache.catalina.webresources.JarWarResourceSet:setManifest() #120        +---[0.010904ms] org.apache.tomcat.util.compat.JreCompat:isJre9Available() #121        +---[0.003307ms] org.apache.catalina.webresources.TomcatJarInputStream:getMetaInfEntry() #133        +---[5.5E-4ms] java.util.jar.JarEntry:getName() #135        +---[6.42E-4ms] java.util.HashMap:put() #135        +---[0.001981ms] org.apache.catalina.webresources.TomcatJarInputStream:getManifestEntry() #137        +---[0.064484ms] org.apache.catalina.webresources.TomcatJarInputStream:close() #141        +---[0.007961ms] org.apache.catalina.webresources.JarWarResourceSet:closeJarFile() #151        `---[0.004643ms] java.io.InputStream:close() #155

發現了一個值得暫停思考的點:

+---[min=0.004452ms,max=34.479307ms,total=74.206249ms,count=31] org.apache.catalina.webresources.TomcatJarInputStream:getNextJarEntry() #117

這行程式碼載入了 31 次,一共耗時 74ms;從名字上看,應該是 tomcat 載入 jar 包時的耗時,那麼是載入了 31 個 jar 包的耗時,還是載入了 jar 包內的某些資源 31 次耗時呢?

TomcatJarInputStream 這個類原始碼的註釋寫到:

The purpose of this sub-class is to obtain references to the JarEntry objectsfor META-INF/ and META-INF/MANIFEST.MF that are otherwise swallowed by theJarInputStream implementation.

大概意思也就是,獲取 jar 包內 META-INF/,META-INF/MANIFEST 的資源,這是一個子類,更多的功能在父類 JarInputStream 裡。

其實看到這裡大概也能猜到問題了,tomcat 載入 jar 包內 META-INF/,META-INF/MANIFEST 的資源導致的耗時,至於為什麼連續請求不會耗時,應該是 tomcat 的快取機制(下面介紹原始碼分析)。

不著急定位問題,試著透過 Arthas 最終定位問題細節,繼續手動深入 trace。

[arthas@24851]$ trace org.apache.catalina.webresources.TomcatJarInputStream *Press Q or Ctrl+C to abort.Affect(class-cnt:1 , method-cnt:4) cost in 44 ms.`---ts=2019-09-14 21:37:47;thread_name=http-nio-7744-exec-5;id=14;is_daemon=true;priority=5;TCCL=org.springframework.boot.loader.LaunchedURLClassLoader@20ad9418    `---[0.234952ms] org.apache.catalina.webresources.TomcatJarInputStream:createZipEntry()        +---[0.039455ms] java.util.jar.JarInputStream:createZipEntry() #43        `---[0.007827ms] java.lang.String:equals() #44`---ts=2019-09-14 21:37:47;thread_name=http-nio-7744-exec-5;id=14;is_daemon=true;priority=5;TCCL=org.springframework.boot.loader.LaunchedURLClassLoader@20ad9418    `---[0.050222ms] org.apache.catalina.webresources.TomcatJarInputStream:createZipEntry()        +---[0.001889ms] java.util.jar.JarInputStream:createZipEntry() #43        `---[0.001643ms] java.lang.String:equals() #46#這裡一共31個trace日誌,刪減了剩下的

從方法名上看,還是載入資源之類的意思。都已經到 jdk 原始碼了,這時候來看一下 TomcatJarInputStream 這個類的原始碼:

/** * Creates a new <code>JarEntry</code> (<code>ZipEntry</code>) for the * specified JAR file entry name. The manifest attributes of * the specified JAR file entry name will be copied to the new * <CODE>JarEntry</CODE>. * * @param name the name of the JAR/ZIP file entry * @return the <code>JarEntry</code> object just created */protected ZipEntry createZipEntry(String name) {    JarEntry e = new JarEntry(name);    if (man != null) {        e.attr = man.getAttributes(name);    }    return e;}

這個 createZipEntry 有個 name 引數,從註釋上看,是 jar/zip 檔名,如果能得到檔名這種關鍵資訊,就可以直接定位問題了;還是透過 Arthas,使用watch 命令,動態監測方法呼叫資料。

watch 方法執行資料觀測

讓你能方便的觀察到指定方法的呼叫情況。能觀察到的範圍為:返回值、丟擲異常、入參,透過編寫 OGNL 表示式進行對應變數的檢視。

watch 該方法的入參:

[arthas@24851]$ watch  org.apache.catalina.webresources.TomcatJarInputStream createZipEntry "{params[0]}"Press Q or Ctrl+C to abort.Affect(class-cnt:1 , method-cnt:1) cost in 27 ms.ts=2019-09-14 21:51:14; [cost=0.14547ms] result=@ArrayList[    @String[META-INF/],]ts=2019-09-14 21:51:14; [cost=0.048028ms] result=@ArrayList[    @String[META-INF/MANIFEST.MF],]ts=2019-09-14 21:51:14; [cost=0.046071ms] result=@ArrayList[    @String[META-INF/resources/],]ts=2019-09-14 21:51:14; [cost=0.033855ms] result=@ArrayList[    @String[META-INF/resources/swagger-ui.html],]ts=2019-09-14 21:51:14; [cost=0.039138ms] result=@ArrayList[    @String[META-INF/resources/webjars/],]ts=2019-09-14 21:51:14; [cost=0.033701ms] result=@ArrayList[    @String[META-INF/resources/webjars/springfox-swagger-ui/],]ts=2019-09-14 21:51:14; [cost=0.033644ms] result=@ArrayList[    @String[META-INF/resources/webjars/springfox-swagger-ui/favicon-16x16.png],]ts=2019-09-14 21:51:14; [cost=0.033976ms] result=@ArrayList[    @String[META-INF/resources/webjars/springfox-swagger-ui/springfox.css],]ts=2019-09-14 21:51:14; [cost=0.032818ms] result=@ArrayList[    @String[META-INF/resources/webjars/springfox-swagger-ui/swagger-ui-standalone-preset.js.map],]ts=2019-09-14 21:51:14; [cost=0.04651ms] result=@ArrayList[    @String[META-INF/resources/webjars/springfox-swagger-ui/swagger-ui.css],]ts=2019-09-14 21:51:14; [cost=0.034793ms] result=@ArrayList[    @String[META-INF/resources/webjars/springfox-swagger-ui/swagger-ui.js.map],

這下直接看到了具體載入的資源名,這麼熟悉的名字:swagger-ui,一個國外的 rest 介面文件工具,又有國內開發者基於 swagger-ui 做了一套 spring mvc 的整合工具,透過註解就可以自動生成 swagger-ui 需要的介面定義 json 檔案,用起來還比較方便,就是侵入性較強。

其實這是 tomcat-embed 的一個 bug 吧,下面詳細介紹一下該 Bug。

Tomcat embed Bug 分析 & 解決

原始碼分析過程實在太漫長,而且也不是本文的重點,所以就不介紹了, 下面直接介紹下分析結果。

順便貼一張 tomcat 處理請求的核心類圖:

1. 為什麼每次請求會載入 Jar 包內的靜態資源?

關鍵在於 org.apache.catalina.mapper.Mapper#internalMapWrapper 這個方法,該版本下處理請求的方式有問題,導致每次都校驗靜態資源。

2. 為什麼連續請求不會出現問題?

因為 Tomcat 對於這種靜態資源的解析是有快取的,優先從快取查詢,快取過期後再重新解析。具體參考 org.apache.catalina.webresources.Cache,預設過期時間 ttl 是 5000ms。

3. 為什麼本地不會復現?

其實確切的說,是透過 spring-boot 打包外掛後不能復現。由於啟動方式的不同,tomcat 使用了不同的類去處理靜態資源,所以沒問題。

4. 如何解決?1)升級 tomcat-embed 版本即可

當前出現 Bug 的版本為:spring-boot:2.0.2.RELEASE,內建的 tomcat embed 版本為 8.5.31。升級 tomcat embed 版本至 8.5.40+ 即可解決此問題,新版本已經修復了。

2)透過替換 springboot pom properties 方式

如果專案是 maven 是繼承的 springboot,即 parent 配置為 springboot 的,或者 dependencyManagement 中 import spring boot 包的。

<properties>    <tomcat.version>8.5.40</tomcat.version></properties>
3)升級 spring boot 版本

springboot 2.1.0.RELEASE 中的 tomcat embed 版本已經大於 8.5.31 了,所以直接將 springboot 升級至該版本及以上版本就可以解決此問題。

作者簡介

空無,Arthas 鐵粉,一個熱愛技術熱愛分享的程式設計師,專注 JAVA 後端開發。

6
最新評論
  • BSA-TRITC(10mg/ml) TRITC-BSA 牛血清白蛋白改性標記羅丹明
  • 跪求!阿里P9手寫的這份1530頁的Java核心程式設計技術手冊